Page 81 - 《软件学报》2026年第2期
P. 81

560                                                        软件学报  2026  年第  37  卷第  2  期


                 异过程中需要的变异位置和名词词组来提升生成测试输入的自然性. 最后, QALT                         通过基于依存句法分析的选择
                 策略提升测试方法的精准度. 实验结果表明, 与最前沿的测试方法相比, QALT                      在两个问答系统中预期比          QAQA
                 和  QAAskeR  分别多检测  2 141  和  4 867  个真阳性缺陷. 消融实验证明了    QALT  的  3  个核心组件均对整体方法起到
                 正向作用. 此外, 本文使用      QALT  生成的测试输入对待测模型进行微调以修复缺陷. 微调后模型实现了                      29.32%  的
                 缺陷修复比率.


                 References
                  [1]   Zhang ZS, Zhao H, Wang R. Machine reading comprehension: The role of contextualized language models and beyond. arXiv:2005.
                     06249, 2020.
                  [2]   Hirschman  L,  Gaizauskas  R.  Natural  language  question  answering:  The  view  from  here.  Natural  Language  Engineering,  2001,  7(4):
                     275–300. [doi: 10.1017/S1351324901002807]
                  [3]   Xiong  CM,  Zhong  V,  Socher  R.  Dynamic  coattention  networks  for  question  answering.  In:  Proc.  of  the  5th  Int’l  Conf.  on  Learning
                     Representations. Toulon: OpenReview.net, 2017.
                  [4]   Clark C, Lee K, Chang MW, Kwiatkowski T, Collins M, Toutanova K. BoolQ: Exploring the surprising difficulty of natural yes/no
                     questions. In: Proc. of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language
                     Technologies. Minneapolis: ACL, 2019. 2924–2936. [doi: 10.18653/v1/N19-1300]
                  [5]   DeLong B. AlphaChat: Underappreciated moments in economic history. 2016. https://equitablegrowth.org/alphachat-underappreciated-
                     moments-in-economic-history/
                  [6]   Kaplan A, Haenlein M. Siri, Siri, in my hand: Who’s the fairest in the land? On the interpretations, illustrations, and implications of
                     artificial intelligence. Business Horizons, 2019, 62(1): 15–25. [doi: 10.1016/j.bushor.2018.08.004]
                  [7]   Roumeliotis KI, Tselikas ND. ChatGPT and Open-AI models: A preliminary review. Future Internet, 2023, 15(6): 192. [doi: 10.3390/
                     fi15060192]
                  [8]   Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C L, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton
                     F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R. Training language models to follow instructions with
                     human feedback. In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: ACM, 2022. 2011.
                  [9]   Zhong WK, Ge JD, Chen X, Li CY, Tang Z, Luo B. Multi-granularity metamorphic testing for neural machine translation system. Ruan
                     Jian Xue Bao/Journal of Software, 2021, 32(4): 1051–1066 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/6221.
                     htm [doi: 10.13328/j.cnki.jos.006221]
                 [10]   Sun ZY, Zhang JM, Harman M, Papadakis M, Zhang L. Automatic testing and improvement of machine translation. In: Proc. of the 42nd
                     ACM/IEEE Int’l Conf. on Software Engineering. Seoul: ACM, 2020. 974–985. [doi: 10.1145/3377811.3380420]
                 [11]   Li Z, Pan MX, Zhang T, Li XD. Testing DNN-based autonomous driving systems under critical environmental conditions. In: Proc. of the
                     38th Int’l Conf. on Machine Learning. PMLR, 2021. 6471–6482.
                 [12]   Eger S, Benz Y. From hero to zéroe: A benchmark of low-level adversarial attacks. In: Proc. of the 1st Conf. of the Asia-Pacific Chapter
                     of the Association for Computational Linguistics and the 10th Int’l Joint Conf. on Natural Language Processing. Suzhou: ACL, 2020.
                     786–803. [doi: 10.18653/v1/2020.aacl-main.79]
                 [13]   Longpre S, Perisetla K, Chen A, Ramesh N, DuBois C, Singh S. Entity-based knowledge conflicts in question answering. In: Proc. of the
                     2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 7052–7063. [doi: 10.18653/v1/2021.emnlp-
                     main.565]
                 [14]   Jia R, Liang P. Adversarial examples for evaluating reading comprehension systems. In: Proc. of the 2017 Conf. on Empirical Methods in
                     Natural Language Processing. Copenhagen: ACL, 2017. 2021–2031. [doi: 10.18653/v1/D17-1215]
                 [15]   Ribeiro M T, Wu TS, Guestrin C, Sameer Singh S. Beyond accuracy: Behavioral testing of NLP models with CheckList. In: Proc. of the
                     58th Annual Meeting of the Association for Computational Linguistics. ACL, 2020. 4902–4912. [doi: 10.18653/v1/2020.acl-main.442]
                 [16]   Chen SQ, Jin S, Xie XY. Testing your question answering software via asking recursively. In: Proc. of the 36th IEEE/ACM Int’l Conf. on
                     Automated Software Engineering (ASE). Melbourne: IEEE, 2021. 104–116. [doi: 10.1109/ASE51524.2021.9678670]
                 [17]   Shen QC, Chen JJ, Zhang JM, Wang HY, Liu S, Tian MH. Natural test generation for precise testing of question answering software. In:
                     Proc.  of  the  37th  IEEE/ACM  Int’l  Conf.  on  Automated  Software  Engineering.  Rochester:  ACM,  2022.  71.  [doi:  10.1145/3551349.
                     3556953]
                 [18]   Team GLM. ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools. arXiv:2406.12793, 2024.
                 [19]   Rajpurkar P, Zhang J, Lopyrev K, Liang P. SQuAD: 100,000+ questions for machine comprehension of text. In: Proc. of the 2016 Conf.
   76   77   78   79   80   81   82   83   84   85   86