Page 82 - 《软件学报》2026年第2期
P. 82
沈庆超 等: 智能问答系统逻辑推理测试 561
on Empirical Methods in Natural Language Processing. Austin: ACL, 2016. 2383–2392. [doi: 10.18653/v1/D16-1264]
[20] Kočiský T, Schwarz J, Blunsom P, Dyer C, Hermann K M, Melis G, Grefenstette E. The NarrativeQA reading comprehension challenge.
Trans. of the Association for Computational Linguistics, 2018, 6: 317–328. [doi: 10.1162/tacl_a_00023]
[21] Sharma V, Kalra A, Vaibhav, Chaudhary S, Patel L, Morency LP. Attend and attack: Attention guided adversarial attacks on visual
question answering models. In: Proc. of the 32nd Conf. on Neural Information Processing Systems. Montreal: NeurIPS, 2018.
[22] Tang RX, Ma C, Zhang WE, Wu Q, Yang XK. Semantic equivalent adversarial data augmentation for visual question answering. In:
Proc. of the 16th European Conf. on Computer Vision. Glasgow: Springer, 2020. 437–453. [doi: 10.1007/978-3-030-58529-7_26]
[23] Sheng SS, Singh A, Goswami V, Magana JAL, Thrush T, Galuba W, Parikh D, Kiela D. Human-adversarial visual question answering.
In: Proc. of the 35th Int’l Conf. on Neural Information Processing Systems. ACM, 2021. 1556.
[24] Walmer M, Sikka K, Sur I, Shrivastava A, Jha S. Dual-key multimodal backdoors for visual question answering. In: Proc. of the 2022
IEEE/CVF Conf. on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022. 15354–15364. [doi: 10.1109/CVPR52688.
2022.01494]
[25] Khashabi D, Khot T, Sabharwal A. More bang for your buck: Natural perturbation for robust question answering. In: Proc. of the 2020
Conf. on Empirical Methods in Natural Language Processing. ACL, 2020. 163–170. [doi: 10.18653/v1/2020.emnlp-main.12]
[26] Khashabi D, Chaturvedi S, Roth M, Upadhyay S, Roth D. Looking beyond the surface: A challenge set for reading comprehension over
multiple sentences. In: Proc. of the 2018 Conf. of the North American Chapter of the Association for Computational Linguistics: Human
Language Technologies. New Orleans: ACL, 2018. 252–262. [doi: 10.18653/v1/N18-1023]
[27] Rajpurkar P, Jia R, Liang P. Know what you don’t know: Unanswerable questions for SQuAD. In: Proc. of the 56th Annual Meeting of
the Association for Computational Linguistics. Melbourne: ACL, 2018. 784–789. [doi: 10.18653/v1/P18-2124]
[28] Trischler A, Wang T, Yuan XD, Harris J, Sordoni A, Bachman P, Suleman K. NewsQA: A machine comprehension dataset. In: Proc. of
the 2nd Workshop on Representation Learning for NLP. Vancouver: ACL, 2017. 191–200. [doi: 10.18653/v1/W17-2623]
[29] Dasigi P, Liu NF, Marasović A, Smith NA, Gardner M. Quoref: A reading comprehension dataset with questions requiring coreferential
reasoning. In: Proc. of the 2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural
Language Processing. Hong Kong: ACL, 2019. 5925–5932. [doi: 10.18653/v1/D19-1606]
[30] Kwiatkowski T, Palomaki J, Redfield O, Collins M, Parikh A, Alberti C, Epstein D, Polosukhin I, Devlin J, Lee K, Toutanova K, Jones L,
Kelcey M, Chang MW, Dai AM, Uszkoreit J, Le Q, Petrov S. Natural questions: A benchmark for question answering research. Trans. of
the Association for Computational Linguistics, 2019, 7: 452–466. [doi: 10.1162/tacl_a_00276]
[31] Dua D, Wang YZ, Dasigi P, Stanovsky G, Singh S, Gardner M. DROP: A reading comprehension benchmark requiring discrete reasoning
over paragraphs. In: Proc. of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human
Language Technologies. Minneapolis: ACL, 2019. 2368–2378. [doi: 10.18653/v1/N19-1246]
[32] Tu KY, Jiang MY, Ding ZH. A metamorphic testing approach for assessing question answering systems. Mathematics, 2021, 9(7): 726.
[doi: 10.3390/math9070726]
[33] Gupta M, Kulkarni N, Chanda R, Rayasam A, Lipton ZC. AmazonQA: A review-based question answering task. In: Proc. of the 28th Int’l
Joint Conf. on Artificial Intelligence. Macao: IJCAI, 2019. 4996–5002. [doi: 10.24963/ijcai.2019/694]
[34] Talmor A, Berant J. MultiQA: An empirical investigation of generalization and transfer in reading comprehension. In: Proc. of the 57th
Annual Meeting of the Association for Computational Linguistics. Florence: ACL, 2019. 4911–4921. [doi: 10.18653/v1/P19-1485]
[35] Clark C, Gardner M. Simple and effective multi-paragraph reading comprehension. In: Proc. of the 56th Annual Meeting of the
Association for Computational Linguistics. Melbourne: ACL, 2018. 845–855. [doi: 10.18653/v1/P18-1078]
[36] Khashabi D, Min S, Khot T, Sabharwal A, Tafjord O, Clark P, Hajishirzi H. UnifiedQA: Crossing format boundaries with a single QA
system. In: Proc. of the 2020 Findings of the Association for Computational Linguistics. ACL, 2020. 1896–1907. [doi: 10.18653/v1/2020.
findings-emnlp.171]
[37] Du ZX, Qian YJ, Liu X, Ding M, Qiu JZ, Yang ZL, Tang J. GLM: General language model pretraining with autoregressive blank
infilling. In: Proc. of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin: ACL, 2022. 320–335. [doi: 10.
18653/v1/2022.acl-long.26]
[38] Chen TY, Cheung SC, Yiu SM. Metamorphic testing: A new approach for generating next test cases. arXiv:2002.12543, 2020.
[39] Huang JT, Zhang JP, Wang WX, He PJ, Su YX, Lyu MR. AEON: A method for automatic evaluation of NLP test cases. In: Proc. of the
31st ACM SIGSOFT Int’l Symp. on Software Testing and Analysis. ACM, 2020. 202–214. [doi: 10.1145/3533767.3534394]
[40] Reimers N, Gurevych I. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In: Proc. of the 2019 Conf. on Empirical
Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural Language Processing. Hong Kong: ACL, 2019.
3982–3992. [doi: 10.18653/v1/D19-1410]

