Page 38 - 《软件学报》2026年第5期
P. 38
韦舒羽 等: 基于音频-语言模型的端到端说话人日志系统 1917
步展示了“说话人损失”对两个训练目标的权衡, 以及 GRPO 阶段起到的突破监督学习瓶颈的关键作用.
然而, 在复杂会议录音和跨场景测试中, 加权微调相比朴素的监督学习反而出现性能退化, 揭示了模型在应对
复杂声学场景时的不足, 以及其对训练-测试数据分布偏移的敏感性. 本文讨论了可能的改进方向, 包括模型压缩
与量化以降低资源消耗、使用支持流式输入的音频编码器扩展输入音频时长限制、采用课程学习和拒绝采样等
方法提升训练稳定性与跨域泛化等.
总之, 本研究为音频-语言模型在说话人日志任务上的应用提供了系统化思路和实践经验, 验证了方案的可行
性, 同时也指出了未来需突破的技术难点, 为后续研究提供了参考.
References
[1] Rubenstein PK, Asawaroengchai C, Nguyen DD, et al. AudioPaLM: A large language model that can speak and listen. arXiv:2306.12925,
2023.
[2] Chu YF, Xu J, Zhou XH, Yang Q, Zhang SL, Yan ZJ, Zhou C, Zhou JR. Qwen-Audio: Advancing universal audio understanding via
unified large-scale audio-language models. arXiv:2311.07919, 2023.
[3] Step-Audio Team. Step-Audio: Unified understanding and generation in intelligent speech interaction. arXiv:2502.11946, 2025.
[4] Peng J, Wang YC, Li BH, Guo YW, Wang HK, Fang YG, Xi Y, Li HY, Li X, Zhang K, Wang S, Yu K. A survey on speech large
language models for understanding. arXiv:2410.18908, 2025.
[5] Chu YF, Xu J, Yang Q, Wei HJ, Wei XP, Guo ZF, Leng YC, Lv YJ, He JZ, Lin JY, Zhou C, Zhou JR. Qwen2-Audio technical report.
arXiv:2407.10759, 2024.
[6] Shao ZH, Wang PY, Zhu QH, Xu RX, Song JX, Bi X, Zhang HW, Zhang MC, Li YK, Wu Y, Guo DY. DeepSeekMath: Pushing the
limits of mathematical reasoning in open language models. arXiv:2402.03300, 2024.
[7] von Neumann T, Boeddeker C, Cord-Landwehr T, Delcroix M, Haeb-Umbach R. Meeting recognition with continuous speech separation
and transcription-supported diarization. In: Proc. of the 2024 IEEE Int’l Conf. on Acoustics, Speech, and Signal Processing Workshops.
Seoul: IEEE, 2024. 775–779. [doi: 10.1109/ICASSPW62465.2024.10625894]
[8] Morrone G, Cornell S, Raj D, Serafini L, Zovato E, Brutti A, Squartini S. Low-latency speech separation guided diarization for telephone
conversations. In: Proc. of the 2022 IEEE Spoken Language Technology Workshop (SLT). Doha: IEEE, 2022. 641–646. [doi: 10.1109/
SLT54892.2023.10023280]
[9] Gruttadauria E, Fontaine M, Essid S. Online speaker diarization of meetings guided by speech separation. In: Proc. of the 2024 IEEE Int’l
Conf. on Acoustics, Speech and Signal Processing (ICASSP). Seoul: IEEE, 2024. 11356–11360. [doi: 10.1109/ICASSP48485.2024.
10447682]
[10] Kalda J, Pagés C, Marxer R, Alumäe T, Bredin H. PixIT: Joint training of speaker diarization and speech separation from real-world multi-
speaker recordings. In: Proc. of the 2024 Speaker and Language Recognition Workshop (Odyssey 2024). Quebec City, 2024. 115–122.
[doi: 10.21437/odyssey.2024-17]
[11] Landini F, Profant J, Diez M, Burget L. Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: Theory,
implementation and analysis on standard tasks. Computer Speech & Language, 2022, 71: 101254. [doi: 10.1016/j.csl.2021.101254]
[12] Zhu BS, Mao QR, Gao LJ, Shen YX. Temporal-segment-and-regroup clustering for speaker diarization. Application Research of
Computers, 2024, 41(9): 2649–2654 (in Chinese with English abstract). [doi: 10.19734/j.issn.1001-3695.2024.01.0017]
[13] Mao QQ, Jia HJ, Zhu BS. Multi-prototype driven graph neural network for speaker diarization. Application Research of Computers,
2025, 42(6): 1778–1783 (in Chinese with English abstract). [doi: 10.19734/j.issn.1001-3695.2024.11.0458]
[14] Kanda N, Xiao X, Gaur Y, Wang XF, Meng Z, Chen Z, Yoshioka T. Transcribe-to-diarize: Neural speaker diarization for unlimited
number of speakers using end-to-end speaker-attributed ASR. In: Proc. of the 2022 IEEE Int’l Conf. on Acoustics, Speech and Signal
Processing (ICASSP). Singapore: IEEE, 2022. 8082–8086. [doi: 10.1109/ICASSP43922.2022.9746225]
[15] Li YZ, Yu F, Liang YH, Guo PC, Shi MH, Du ZH, Zhang SL, Xie L. Sa-Paraformer: Non-autoregressive end-to-end speaker-attributed
ASR. In: Proc. of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). Taipei: IEEE, 2023. 1–7. [doi:
10.1109/ASRU57964.2023.10389762]
[16] Kanda N, Ye GL, Gaur Y, Wang XF, Meng Z, Chen Z, Yoshioka T. End-to-end speaker-attributed ASR with Transformer. arXiv:
2104.02128, 2021.
[17] Wang Q, Huang YL, Zhao GL, Clark E, Xia W, Liao H. DiarizationLM: Speaker diarization post-processing with large language models.
In: Proc. of the 25th Annual Conf. of the Int’l Speech Communication Association. Kos, 2024. 3754–3758. [doi: 10.21437/Interspeech.
2024-209]
[18] Efstathiadis G, Yadav V, Abbas A. LLM-based speaker diarization correction: A generalizable approach. Speech Communication, 2025,

