Page 316 - 《软件学报》2026年第4期
P. 316
王骞玥 等: 大语言模型驱动的可信政务问答技术 1757
[10] Fang KY, Xu KW. Automating government response to citizens’ questions: A large language model-based question-answering guidance
generation system. In: Proc. of the 3rd Int’l Conf. on Digital Society and Intelligent Systems. Chengdu: IEEE, 2023. 386–389. [doi: 10.
1109/DSInS60115.2023.10455136]
[11] Liu Y, Yao YS, Ton JF, Zhang XY, Guo RC, Cheng H, Klochkov Y, Taufiq MF, Li H. Trustworthy LLMs: A survey and guideline for
evaluating large language models’ alignment. arXiv:2308.05374, 2024.
[12] Dhuliawala S, Komeili M, Xu J, Raileanu R, Li X, Celikyilmaz A, Weston J. Chain-of-verification reduces hallucination in large
language models. In: Findings of the Association for Computational Linguistics: ACL 2024. Bangkok: ACL, 2024. 3563–3578. [doi: 10.
18653/v1/2024.findings-acl.212]
[13] Wei J, Wang XZ, Schuurmans D, Bosma M, Ichter B, Xia F, Chi EH, Le QV, Zhou D. Chain-of-thought prompting elicits reasoning in
large language models. In: Proc. of the 36th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc.,
2022. 1800.
[14] Wen YL, Wang ZF, Sun JM. MindMap: Knowledge graph prompting sparks graph of thoughts in large language models. In: Proc. of the
62nd Annual Meeting of the Association for Computational Linguistics. Bangkok: ACL, 2024. 10370–10388. [doi: 10.18653/v1/2024.acl-
long.558]
[15] Jiang ZB, Araki J, Ding HB, Neubig G. How can we know when language models know? On the calibration of language models for
question answering. Trans. of the Association for Computational Linguistics, 2021, 9: 962–977. [doi: 10.1162/tacl_a_00407]
[16] Zhou CT, Liu PF, Xu PX, Iyer S, Sun J, Mao YN, Ma XZ, Efrat A, Yu P, Yu LL, Zhang SS, Ghosh G, Lewis M, Zettlemoyer L, Levy O.
LIMA: Less is more for alignment. In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran
Associates Inc., 2023. 2400.
[17] Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, Küttler H, Lewis M, Yih WT, Rocktäschel T, Riedel S, Kiela D. Retrieval-
augmented generation for knowledge-intensive NLP tasks. In: Proc. of the 34th Int’l Conf. on Neural Information Processing Systems.
Vancouver: Curran Associates Inc., 2020. 793.
[18] Guu K, Lee K, Tung Z, Pasupat P, Chang MW. REALM: Retrieval-augmented language model pre-training. In: Proc. of the 37th Int’l
Conf. on Machine Learning. arXiv:2002.08909v1, 2020.
[19] Zhang SH, Song YL, Yang JH, Li YQ, Han B, Tan MK. Detecting machine-generated texts by multi-population aware optimization for
maximum mean discrepancy. In: Proc. of the 12th Int’l Conf. on Learning Representations. Vienna: OpenReview.net, 2024.
[20] Ma JK, Wang YS, Li G, Mei H. Influence of acquirer’s participation on project performance: An analysis on e-government projects. Ruan
Jian Xue Bao/Journal of Software, 2012, 23(10): 2679–2694 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/4244.
htm [doi: 10.3724/SP.J.1001.2012.04244]
[21] Zeng AH, Liu X, Du ZX, Wang ZH, Lai HY, Ding M, Yang ZY, Xu YF, Zheng WD, Xia X, Tam WL, Ma ZX, Xue YF, Zhai JD, Chen
WG, Liu ZY, Zhang P, Dong YX, Tang J. GLM-130B: An open bilingual pre-trained model. In: Proc. of the 11th Int’l Conf. on Learning
Representations. Kigali: OpenReview.net, 2023. 320–335.
[22] Hu EJ, Shen YL, Wallis P, Allen-Zhu Z, Li YZ, Wang SA, Wang L, Chen WZ. LoRA: Low-rank adaptation of large language models. In:
Proc. of the 10th Int’l Conf. on Learning Representations. OpenReview.net, 2022.
[23] Tian XH. Algorithm integration, information-driven empowerment, and capacity building in grassroots government governance. CASS
Journal of Political Science, 2025(1): 119–134, 190 (in Chinese with English abstract). [doi: 10.3969/j.issn.1000-3355.2025.1.zzxyj
202501010]
[24] Nizon-Deladoeuille M, Stefánsson B, Neukirchen H, Welsh T. Towards supporting penetration testing education with large language
models: An evaluation and comparison. In: Proc. of the 11th Int’l Conf. on Social Networks Analysis, Management and Security. Gran
Canaria: IEEE, 2024. 227–229. [doi: 10.1109/SNAMS64316.2024.10883774]
[25] Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, Leike J, Schulman J, Sutskever I, Cobbe K. Let’s verify step by step. In:
Proc. of the 12th Int’l Conf. on Learning Representations. Vienna: OpenReview.net, 2024.
[26] Bai YS, Lv X, Zhang JJ, Lyu HC, Tang JK, Huang ZD, Du ZX, Liu X, Zeng AH, Hou L, Dong YX, Tang J, Li JZ. LongBench: A
bilingual, multitask benchmark for long context understanding. In: Proc. of the 62nd Annual Meeting of the Association for
Computational Linguistics. Bangkok: ACL, 2024. 3119–3137. [doi: 10.18653/v1/2024.acl-long.172]
[27] Guha N, Nyarko J, Ho DE, et al. LegalBench: A collaboratively built benchmark for measuring legal reasoning in large language models.
In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: NeurIPS, 2023. 44123–44279.
[28] Lin CY. ROUGE: A package for automatic evaluation of summaries. In: Text Summarization Branches Out. Barcelona: ACL, 2004.
74–81.
[29] Papineni K, Roukos S, Ward T, Zhu WJ. Bleu: A method for automatic evaluation of machine translation. In: Proc. of the 40th Annual

