Page 38 - 《软件学报》2026年第5期
P. 38

韦舒羽 等: 基于音频-语言模型的端到端说话人日志系统                                                     1917


                 步展示了“说话人损失”对两个训练目标的权衡, 以及               GRPO  阶段起到的突破监督学习瓶颈的关键作用.
                    然而, 在复杂会议录音和跨场景测试中, 加权微调相比朴素的监督学习反而出现性能退化, 揭示了模型在应对
                 复杂声学场景时的不足, 以及其对训练-测试数据分布偏移的敏感性. 本文讨论了可能的改进方向, 包括模型压缩
                 与量化以降低资源消耗、使用支持流式输入的音频编码器扩展输入音频时长限制、采用课程学习和拒绝采样等
                 方法提升训练稳定性与跨域泛化等.
                    总之, 本研究为音频-语言模型在说话人日志任务上的应用提供了系统化思路和实践经验, 验证了方案的可行
                 性, 同时也指出了未来需突破的技术难点, 为后续研究提供了参考.

                 References
                  [1]   Rubenstein PK, Asawaroengchai C, Nguyen DD, et al. AudioPaLM: A large language model that can speak and listen. arXiv:2306.12925,
                     2023.
                  [2]   Chu YF, Xu J, Zhou XH, Yang Q, Zhang SL, Yan ZJ, Zhou C, Zhou JR. Qwen-Audio: Advancing universal audio understanding via
                     unified large-scale audio-language models. arXiv:2311.07919, 2023.
                  [3]   Step-Audio Team. Step-Audio: Unified understanding and generation in intelligent speech interaction. arXiv:2502.11946, 2025.
                  [4]   Peng J, Wang YC, Li BH, Guo YW, Wang HK, Fang YG, Xi Y, Li HY, Li X, Zhang K, Wang S, Yu K. A survey on speech large
                     language models for understanding. arXiv:2410.18908, 2025.
                  [5]   Chu YF, Xu J, Yang Q, Wei HJ, Wei XP, Guo ZF, Leng YC, Lv YJ, He JZ, Lin JY, Zhou C, Zhou JR. Qwen2-Audio technical report.
                     arXiv:2407.10759, 2024.
                  [6]   Shao ZH, Wang PY, Zhu QH, Xu RX, Song JX, Bi X, Zhang HW, Zhang MC, Li YK, Wu Y, Guo DY. DeepSeekMath: Pushing the
                     limits of mathematical reasoning in open language models. arXiv:2402.03300, 2024.
                  [7]   von Neumann T, Boeddeker C, Cord-Landwehr T, Delcroix M, Haeb-Umbach R. Meeting recognition with continuous speech separation
                     and transcription-supported diarization. In: Proc. of the 2024 IEEE Int’l Conf. on Acoustics, Speech, and Signal Processing Workshops.
                     Seoul: IEEE, 2024. 775–779. [doi: 10.1109/ICASSPW62465.2024.10625894]
                  [8]   Morrone G, Cornell S, Raj D, Serafini L, Zovato E, Brutti A, Squartini S. Low-latency speech separation guided diarization for telephone
                     conversations. In: Proc. of the 2022 IEEE Spoken Language Technology Workshop (SLT). Doha: IEEE, 2022. 641–646. [doi: 10.1109/
                     SLT54892.2023.10023280]
                  [9]   Gruttadauria E, Fontaine M, Essid S. Online speaker diarization of meetings guided by speech separation. In: Proc. of the 2024 IEEE Int’l
                     Conf.  on  Acoustics,  Speech  and  Signal  Processing  (ICASSP).  Seoul:  IEEE,  2024.  11356–11360.  [doi:  10.1109/ICASSP48485.2024.
                     10447682]
                 [10]   Kalda J, Pagés C, Marxer R, Alumäe T, Bredin H. PixIT: Joint training of speaker diarization and speech separation from real-world multi-
                     speaker recordings. In: Proc. of the 2024 Speaker and Language Recognition Workshop (Odyssey 2024). Quebec City, 2024. 115–122.
                     [doi: 10.21437/odyssey.2024-17]
                 [11]   Landini  F,  Profant  J,  Diez  M,  Burget  L.  Bayesian  HMM  clustering  of  x-vector  sequences  (VBx)  in  speaker  diarization:  Theory,
                     implementation and analysis on standard tasks. Computer Speech & Language, 2022, 71: 101254. [doi: 10.1016/j.csl.2021.101254]
                 [12]   Zhu  BS,  Mao  QR,  Gao  LJ,  Shen  YX.  Temporal-segment-and-regroup  clustering  for  speaker  diarization.  Application  Research  of
                     Computers, 2024, 41(9): 2649–2654 (in Chinese with English abstract). [doi: 10.19734/j.issn.1001-3695.2024.01.0017]
                 [13]   Mao QQ, Jia HJ, Zhu BS. Multi-prototype driven graph neural network for speaker diarization. Application Research of Computers,
                     2025, 42(6): 1778–1783 (in Chinese with English abstract). [doi: 10.19734/j.issn.1001-3695.2024.11.0458]
                 [14]   Kanda N, Xiao X, Gaur Y, Wang XF, Meng Z, Chen Z, Yoshioka T. Transcribe-to-diarize: Neural speaker diarization for unlimited
                     number of speakers using end-to-end speaker-attributed ASR. In: Proc. of the 2022 IEEE Int’l Conf. on Acoustics, Speech and Signal
                     Processing (ICASSP). Singapore: IEEE, 2022. 8082–8086. [doi: 10.1109/ICASSP43922.2022.9746225]
                 [15]   Li YZ, Yu F, Liang YH, Guo PC, Shi MH, Du ZH, Zhang SL, Xie L. Sa-Paraformer: Non-autoregressive end-to-end speaker-attributed
                     ASR. In: Proc. of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). Taipei: IEEE, 2023. 1–7. [doi:
                     10.1109/ASRU57964.2023.10389762]
                 [16]   Kanda  N,  Ye  GL,  Gaur  Y,  Wang  XF,  Meng  Z,  Chen  Z,  Yoshioka  T.  End-to-end  speaker-attributed  ASR  with  Transformer.  arXiv:
                     2104.02128, 2021.
                 [17]   Wang Q, Huang YL, Zhao GL, Clark E, Xia W, Liao H. DiarizationLM: Speaker diarization post-processing with large language models.
                     In: Proc. of the 25th Annual Conf. of the Int’l Speech Communication Association. Kos, 2024. 3754–3758. [doi: 10.21437/Interspeech.
                     2024-209]
                 [18]   Efstathiadis G, Yadav V, Abbas A. LLM-based speaker diarization correction: A generalizable approach. Speech Communication, 2025,
   33   34   35   36   37   38   39   40   41   42   43