Page 81 - 《软件学报》2026年第2期
P. 81
560 软件学报 2026 年第 37 卷第 2 期
异过程中需要的变异位置和名词词组来提升生成测试输入的自然性. 最后, QALT 通过基于依存句法分析的选择
策略提升测试方法的精准度. 实验结果表明, 与最前沿的测试方法相比, QALT 在两个问答系统中预期比 QAQA
和 QAAskeR 分别多检测 2 141 和 4 867 个真阳性缺陷. 消融实验证明了 QALT 的 3 个核心组件均对整体方法起到
正向作用. 此外, 本文使用 QALT 生成的测试输入对待测模型进行微调以修复缺陷. 微调后模型实现了 29.32% 的
缺陷修复比率.
References
[1] Zhang ZS, Zhao H, Wang R. Machine reading comprehension: The role of contextualized language models and beyond. arXiv:2005.
06249, 2020.
[2] Hirschman L, Gaizauskas R. Natural language question answering: The view from here. Natural Language Engineering, 2001, 7(4):
275–300. [doi: 10.1017/S1351324901002807]
[3] Xiong CM, Zhong V, Socher R. Dynamic coattention networks for question answering. In: Proc. of the 5th Int’l Conf. on Learning
Representations. Toulon: OpenReview.net, 2017.
[4] Clark C, Lee K, Chang MW, Kwiatkowski T, Collins M, Toutanova K. BoolQ: Exploring the surprising difficulty of natural yes/no
questions. In: Proc. of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language
Technologies. Minneapolis: ACL, 2019. 2924–2936. [doi: 10.18653/v1/N19-1300]
[5] DeLong B. AlphaChat: Underappreciated moments in economic history. 2016. https://equitablegrowth.org/alphachat-underappreciated-
moments-in-economic-history/
[6] Kaplan A, Haenlein M. Siri, Siri, in my hand: Who’s the fairest in the land? On the interpretations, illustrations, and implications of
artificial intelligence. Business Horizons, 2019, 62(1): 15–25. [doi: 10.1016/j.bushor.2018.08.004]
[7] Roumeliotis KI, Tselikas ND. ChatGPT and Open-AI models: A preliminary review. Future Internet, 2023, 15(6): 192. [doi: 10.3390/
fi15060192]
[8] Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C L, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton
F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R. Training language models to follow instructions with
human feedback. In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: ACM, 2022. 2011.
[9] Zhong WK, Ge JD, Chen X, Li CY, Tang Z, Luo B. Multi-granularity metamorphic testing for neural machine translation system. Ruan
Jian Xue Bao/Journal of Software, 2021, 32(4): 1051–1066 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6221.
htm [doi: 10.13328/j.cnki.jos.006221]
[10] Sun ZY, Zhang JM, Harman M, Papadakis M, Zhang L. Automatic testing and improvement of machine translation. In: Proc. of the 42nd
ACM/IEEE Int’l Conf. on Software Engineering. Seoul: ACM, 2020. 974–985. [doi: 10.1145/3377811.3380420]
[11] Li Z, Pan MX, Zhang T, Li XD. Testing DNN-based autonomous driving systems under critical environmental conditions. In: Proc. of the
38th Int’l Conf. on Machine Learning. PMLR, 2021. 6471–6482.
[12] Eger S, Benz Y. From hero to zéroe: A benchmark of low-level adversarial attacks. In: Proc. of the 1st Conf. of the Asia-Pacific Chapter
of the Association for Computational Linguistics and the 10th Int’l Joint Conf. on Natural Language Processing. Suzhou: ACL, 2020.
786–803. [doi: 10.18653/v1/2020.aacl-main.79]
[13] Longpre S, Perisetla K, Chen A, Ramesh N, DuBois C, Singh S. Entity-based knowledge conflicts in question answering. In: Proc. of the
2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 7052–7063. [doi: 10.18653/v1/2021.emnlp-
main.565]
[14] Jia R, Liang P. Adversarial examples for evaluating reading comprehension systems. In: Proc. of the 2017 Conf. on Empirical Methods in
Natural Language Processing. Copenhagen: ACL, 2017. 2021–2031. [doi: 10.18653/v1/D17-1215]
[15] Ribeiro M T, Wu TS, Guestrin C, Sameer Singh S. Beyond accuracy: Behavioral testing of NLP models with CheckList. In: Proc. of the
58th Annual Meeting of the Association for Computational Linguistics. ACL, 2020. 4902–4912. [doi: 10.18653/v1/2020.acl-main.442]
[16] Chen SQ, Jin S, Xie XY. Testing your question answering software via asking recursively. In: Proc. of the 36th IEEE/ACM Int’l Conf. on
Automated Software Engineering (ASE). Melbourne: IEEE, 2021. 104–116. [doi: 10.1109/ASE51524.2021.9678670]
[17] Shen QC, Chen JJ, Zhang JM, Wang HY, Liu S, Tian MH. Natural test generation for precise testing of question answering software. In:
Proc. of the 37th IEEE/ACM Int’l Conf. on Automated Software Engineering. Rochester: ACM, 2022. 71. [doi: 10.1145/3551349.
3556953]
[18] Team GLM. ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools. arXiv:2406.12793, 2024.
[19] Rajpurkar P, Zhang J, Lopyrev K, Liang P. SQuAD: 100,000+ questions for machine comprehension of text. In: Proc. of the 2016 Conf.

