Page 39 - 《软件学报》2026年第5期
P. 39

1918                                                       软件学报  2026  年第  37  卷第  5  期


                     170: 103224. [doi: 10.1016/j.specom.2025.103224]
                 [19]   Yang  M,  Kanda  N,  Wang  XF,  Chen  JK,  Wang  PD,  Xue  J,  Li  JY,  Yoshioka  T.  DiariST:  Streaming  speech  translation  with  speaker
                     diarization.  In:  Proc.  of  the  2024  IEEE  Int’l  Conf.  on  Acoustics,  Speech  and  Signal  Processing  (ICASSP).  Seoul:  IEEE,  2024.
                     10866–10870. [doi: 10.1109/ICASSP48485.2024.10446050]
                 [20]   Piñeiro-Martín A, García-Mateo C, Docío-Fernández L, López-Pérez MDC, Rehm G. Weighted cross-entropy for low-resource languages
                     in  multilingual  speech  recognition.  In:  Proc.  of  the  25th  Annual  Conf.  of  the  Int’l  Speech  Communication  Association.  Kos,  2024.
                     1235–1239. [doi: 10.21437/Interspeech.2024-734]
                 [21]   Khare A, Han EJ, Yang YG, Stolcke A. ASR-aware end-to-end neural diarization. In: Proc. of the 2022 IEEE Int’l Conf. on Acoustics,
                     Speech and Signal Processing (ICASSP). Singapore: IEEE, 2022. 8092–8096. [doi: 10.1109/ICASSP43922.2022.9746964]
                 [22]   Kumar  S,  Madikeri  S,  Nigmatulina  I,  Villatoro-Tello  E,  Motlicek  P,  Pandia  K,  Dubagunta  SP,  Ganapathiraju  A.  Multitask  speech
                     recognition and speaker change detection for unknown number of speakers. In: Proc. of the 2024 IEEE Int’l Conf. on Acoustics, Speech
                     and Signal Processing (ICASSP). Seoul: IEEE, 2024. 12592–12596. [doi: 10.1109/ICASSP48485.2024.10446130]
                 [23]   Li G, Liu JZ, Dinkel H, Niu YD, Zhang JB, Luan J. Reinforcement learning outperforms supervised fine-tuning: A case study on audio
                     question answering. arXiv:2503.11197, 2025.
                 [24]   Nagpal C, Venugopalan S, Tobin J, Ladewig M, Heller K, Tomanek K. Speech recognition with LLMs adapted to disordered speech
                     using reinforcement learning. In: Proc. of the 2025 IEEE Int’l Conf. on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad:
                     IEEE, 2025. 1–5. [doi: 10.1109/ICASSP49660.2025.10888006]
                 [25]   Rafailov R, Sharma A, Mitchell E, Ermon S, Manning CD, Finn C. Direct preference optimization: Your language model is secretly a
                     reward model. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2023.
                     53728–53741.
                 [26]   Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv:1707.06347, 2017.
                 [27]   Wang SH, Zhang SY, Zhang J, Hu RY, Li XY, Zhang TW, Li JW, Wu F, Wang GY, Hovy E. Reinforcement learning enhanced LLMs: A
                     survey. arXiv:2412.10400, 2025.
                 [28]   Park T, Medennikov I, Dhawan K, Wang WQ, Huang H, Koluguri NR, Puvvada KC, Balam J, Ginsburg B. Sortformer: A novel approach
                     for permutation-resolved speaker supervision in speech-to-text systems. arXiv:2409.06656, 2025.
                 [29]   Gao ZF, Zhang SL, McLoughlin I, Yan ZJ. Paraformer: Fast and accurate parallel Transformer for non-autoregressive end-to-end speech
                     recognition. arXiv:2206.08317, 2023.
                 [30]   Mao YX, Wang Q, Chen C, Qu Y, Ji XY. Offline reinforcement learning with OOD state correction and OOD action suppression. In:
                     Proc. of the 38th Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2025. 93568–93601.
                 [31]   Haider  T,  Roscher  K,  da  Roza  FS,  Günnemann  S.  Out-of-distribution  detection  for  reinforcement  learning  agents  with  probabilistic
                     dynamics  models.  In:  Proc.  of  the  2023  Int’l  Conf.  on  Autonomous  Agents  and  Multiagent  Systems.  London:  Int’l  Foundation  for
                     Autonomous Agents and Multiagent Systems, 2023. 851–859.
                 [32]   Gulati A, Qin J, Chiu CC, Parmar N, Zhang Y, Yu JH, Han W, Wang SB, Zhang ZD, Wu YH, Pang RM. Conformer: Convolution-
                     augmented  Transformer  for  speech  recognition.  In:  Proc.  of  the  21st  Annual  Conf.  of  the  Int’l  Speech  Communication  Association.
                     Shanghai, 2020. 5036–5040.
                 [33]   Zhong SS, Gao SH, Huang ZZ, Wen WS, Zitnik M, Zhou P. MoExtend: Tuning new experts for modality and task extension. In: Proc. of
                     the 62nd Annual Meeting of the Association for Computational Linguistics (Vol. 4: Student Research Workshop). Bangkok: ACL, 2024.
                     494–505.

                 附中文参考文献
                 [12]   朱必松, 毛启容, 高利剑, 沈雅馨. 基于时间分段和重组聚类的说话人日志方法. 计算机应用研究, 2024, 41(9): 2649–2654. [doi:
                     10.19734/j.issn.1001-3695.2024.01.0017]
                 [13]   毛青青, 贾洪杰, 朱必松. 面向说话人日志的多原型驱动图神经网络方法. 计算机应用研究, 2025, 42(6): 1778–1783. [doi: 10.19734/j.
                     issn.1001-3695.2024.11.0458]

                 作者简介
                 韦舒羽, 博士生, 主要研究领域为多模态大模型.
                 丘德来, 硕士, 主要研究领域为多模态大模型, 自然语言处理, 智能问答.
                 刘升平, 博士, CCF  专业会员, 主要研究领域为大语言模型, 自然语言处理, 知识图谱.
                 桑基韬, 博士, 教授, 博士生导师, CCF  高级会员, 主要研究领域为多模态智能, 可信与对齐, 推理模型, AI Agent.
   34   35   36   37   38   39   40   41   42   43   44