Page 82 - 《软件学报》2026年第2期
P. 82

沈庆超 等: 智能问答系统逻辑推理测试                                                              561


                     on Empirical Methods in Natural Language Processing. Austin: ACL, 2016. 2383–2392. [doi: 10.18653/v1/D16-1264]
                 [20]   Kočiský T, Schwarz J, Blunsom P, Dyer C, Hermann K M, Melis G, Grefenstette E. The NarrativeQA reading comprehension challenge.
                     Trans. of the Association for Computational Linguistics, 2018, 6: 317–328. [doi: 10.1162/tacl_a_00023]
                 [21]   Sharma  V,  Kalra  A,  Vaibhav,  Chaudhary  S,  Patel  L,  Morency  LP.  Attend  and  attack:  Attention  guided  adversarial  attacks  on  visual
                     question answering models. In: Proc. of the 32nd Conf. on Neural Information Processing Systems. Montreal: NeurIPS, 2018.
                 [22]   Tang RX, Ma C, Zhang WE, Wu Q, Yang XK. Semantic equivalent adversarial data augmentation for visual question answering. In:
                     Proc. of the 16th European Conf. on Computer Vision. Glasgow: Springer, 2020. 437–453. [doi: 10.1007/978-3-030-58529-7_26]
                 [23]   Sheng SS, Singh A, Goswami V, Magana JAL, Thrush T, Galuba W, Parikh D, Kiela D. Human-adversarial visual question answering.
                     In: Proc. of the 35th Int’l Conf. on Neural Information Processing Systems. ACM, 2021. 1556.
                 [24]   Walmer M, Sikka K, Sur I, Shrivastava A, Jha S. Dual-key multimodal backdoors for visual question answering. In: Proc. of the 2022
                     IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 15354–15364. [doi: 10.1109/CVPR52688.
                     2022.01494]
                 [25]   Khashabi D, Khot T, Sabharwal A. More bang for your buck: Natural perturbation for robust question answering. In: Proc. of the 2020
                     Conf. on Empirical Methods in Natural Language Processing. ACL, 2020. 163–170. [doi: 10.18653/v1/2020.emnlp-main.12]
                 [26]   Khashabi D, Chaturvedi S, Roth M, Upadhyay S, Roth D. Looking beyond the surface: A challenge set for reading comprehension over
                     multiple sentences. In: Proc. of the 2018 Conf. of the North American Chapter of the Association for Computational Linguistics: Human
                     Language Technologies. New Orleans: ACL, 2018. 252–262. [doi: 10.18653/v1/N18-1023]
                 [27]   Rajpurkar P, Jia R, Liang P. Know what you don’t know: Unanswerable questions for SQuAD. In: Proc. of the 56th Annual Meeting of
                     the Association for Computational Linguistics. Melbourne: ACL, 2018. 784–789. [doi: 10.18653/v1/P18-2124]
                 [28]   Trischler A, Wang T, Yuan XD, Harris J, Sordoni A, Bachman P, Suleman K. NewsQA: A machine comprehension dataset. In: Proc. of
                     the 2nd Workshop on Representation Learning for NLP. Vancouver: ACL, 2017. 191–200. [doi: 10.18653/v1/W17-2623]
                 [29]   Dasigi P, Liu NF, Marasović A, Smith NA, Gardner M. Quoref: A reading comprehension dataset with questions requiring coreferential
                     reasoning. In: Proc. of the 2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural
                     Language Processing. Hong Kong: ACL, 2019. 5925–5932. [doi: 10.18653/v1/D19-1606]
                 [30]   Kwiatkowski T, Palomaki J, Redfield O, Collins M, Parikh A, Alberti C, Epstein D, Polosukhin I, Devlin J, Lee K, Toutanova K, Jones L,
                     Kelcey M, Chang MW, Dai AM, Uszkoreit J, Le Q, Petrov S. Natural questions: A benchmark for question answering research. Trans. of
                     the Association for Computational Linguistics, 2019, 7: 452–466. [doi: 10.1162/tacl_a_00276]
                 [31]   Dua D, Wang YZ, Dasigi P, Stanovsky G, Singh S, Gardner M. DROP: A reading comprehension benchmark requiring discrete reasoning
                     over paragraphs. In: Proc. of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human
                     Language Technologies. Minneapolis: ACL, 2019. 2368–2378. [doi: 10.18653/v1/N19-1246]
                 [32]   Tu KY, Jiang MY, Ding ZH. A metamorphic testing approach for assessing question answering systems. Mathematics, 2021, 9(7): 726.
                     [doi: 10.3390/math9070726]
                 [33]   Gupta M, Kulkarni N, Chanda R, Rayasam A, Lipton ZC. AmazonQA: A review-based question answering task. In: Proc. of the 28th Int’l
                     Joint Conf. on Artificial Intelligence. Macao: IJCAI, 2019. 4996–5002. [doi: 10.24963/ijcai.2019/694]
                 [34]   Talmor A, Berant J. MultiQA: An empirical investigation of generalization and transfer in reading comprehension. In: Proc. of the 57th
                     Annual Meeting of the Association for Computational Linguistics. Florence: ACL, 2019. 4911–4921. [doi: 10.18653/v1/P19-1485]
                 [35]   Clark  C,  Gardner  M.  Simple  and  effective  multi-paragraph  reading  comprehension.  In:  Proc.  of  the  56th  Annual  Meeting  of  the
                     Association for Computational Linguistics. Melbourne: ACL, 2018. 845–855. [doi: 10.18653/v1/P18-1078]
                 [36]   Khashabi D, Min S, Khot T, Sabharwal A, Tafjord O, Clark P, Hajishirzi H. UnifiedQA: Crossing format boundaries with a single QA
                     system. In: Proc. of the 2020 Findings of the Association for Computational Linguistics. ACL, 2020. 1896–1907. [doi: 10.18653/v1/2020.
                     findings-emnlp.171]
                 [37]   Du  ZX,  Qian  YJ,  Liu  X,  Ding  M,  Qiu  JZ,  Yang  ZL,  Tang  J.  GLM:  General  language  model  pretraining  with  autoregressive  blank
                     infilling. In: Proc. of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin: ACL, 2022. 320–335. [doi: 10.
                     18653/v1/2022.acl-long.26]
                 [38]   Chen TY, Cheung SC, Yiu SM. Metamorphic testing: A new approach for generating next test cases. arXiv:2002.12543, 2020.
                 [39]   Huang JT, Zhang JP, Wang WX, He PJ, Su YX, Lyu MR. AEON: A method for automatic evaluation of NLP test cases. In: Proc. of the
                     31st ACM SIGSOFT Int’l Symp. on Software Testing and Analysis. ACM, 2020. 202–214. [doi: 10.1145/3533767.3534394]
                 [40]   Reimers N, Gurevych I. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In: Proc. of the 2019 Conf. on Empirical
                     Methods  in  Natural  Language  Processing  and  the  9th  Int’l  Joint  Conf.  on  Natural  Language  Processing.  Hong  Kong:  ACL,  2019.
                     3982–3992. [doi: 10.18653/v1/D19-1410]
   77   78   79   80   81   82   83   84   85   86   87