Page 299 - 《软件学报》2026年第2期
P. 299
778 软件学报 2026 年第 37 卷第 2 期
Grammar Error Count: 该指标使用 LanguageTool 工具对句子进行错误检测, 统计其中的错误数量, 主要包括
主谓不一致、句子结构不完整、拼写错误等. 该指标以错误数量计, 数值越低表示句子质量越高 [84] .
上述指标从结构合法性、流畅性判断和细节检测这 3 个层面客观地验证我们方法在提升目标端句法准确性
方面的有效性, 且在业界评估实践中均为常用手段. 实验结果如表 11 所示. 从表中最后两行所有数据测试的平均
值可见, 在 3 项指标上, SynNAT 生成的译文整体优于 GLAT, 表明源端句法建模在目标句法结构的生成中起到了
积极作用.
表 11 蒸馏的 IWSLT14 DE→EN 测试样本评估
模型 句子 Parse Validity Acceptability Score Grammar Error Count
GLAT the only country country. 1 0.760 4 3
SynNAT the only country in the world. 1 0.948 2 2
GLAT she did for three years. 1 0.859 9 2
SynNAT she did it for three years. 1 0.981 7 2
GLAT all data avg 0.482 4 0.480 4 7.469 0
SynNAT all data avg 0.493 4 0.492 6 7.378 7
5 结 论
本研究重新探讨了依存句法信息增强源端是否能够提升 Fully NAT 的性能. 通过将源句的丰富依存句法结构
显式编码到 Fully NAT 模型中, 我们提出的 SynNAT 模型显著提高了翻译质量, 同时不会产生明显的解码速度损
失. 实验结果表明, 该方法在多个翻译基准测试中表现出色, 显著提升了 Fully NAT 模型的翻译性能, 验证了利用
GCN 对依存句法信息进行编码的有效性和潜力. 该创新方法不仅为 NAT 模型的发展提供了宝贵的新见解, 也为
实现高效且高质量的非自回归翻译奠定了基础.
References:
[1] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I. Attention is all you need. In: Proc. of the
31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
[2] OpenAI. GPT-4 technical report. arXiv:2303.08774, 2024.
[3] Dubey A, Jauhri A, Pandey A, et al. The llama 3 herd of models. arXiv:2407.21783, 2024.
[4] Lee J, Stevens N, Han SC, Song M. A survey of large language models in finance (FinLLMs). arXiv:2402.02315, 2024.
[5] Cai TL, Li YH, Geng ZY, Peng HW, Lee JD, Chen DM, Dao T. Medusa: Simple LLM inference acceleration framework with multiple
decoding heads. arXiv:2401.10774, 2024.
[6] Gu JT, Bradbury J, Xiong CM, Li VOK, Socher R. Non-autoregressive neural machine translation. arXiv:1711.02281, 2018.
[7] Guo P, Xiao YS, Li JT, Ji YX, Zhang M. Isotropy-enhanced conditional masked language models. In: Proc. of the 2023 Conf. on
Empirical Methods in Natural Language Processing. Singapore: ACL, 2023. 8278–8289. [doi: 10.18653/v1/2023.findings-emnlp.555]
[8] Xiao YS, Wu LJ, Guo JL, Li JT, Zhang M, Qin T, Liu TY. A survey on non-autoregressive generation for neural machine translation and
beyond. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2023, 45(10): 11407–11427. [doi: 10.1109/TPAMI.2023.3277122]
[9] Ghazvininejad M, Levy O, Liu YH, Zettlemoyer L. Mask-predict: Parallel decoding of conditional masked language models. In: Proc. of
the 2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural Language Processing
(EMNLP-IJCNLP). Hong Kong: ACL, 2019. 6112–6121. [doi: 10.18653/V1/D19-1633]
[10] Gu JT, Wang CH, Zhao JB. Levenshtein Transformer. In: Proc. of the 33rd Int’l Conf. on Neural Information Processing Systems.
Vancouver: Curran Associates Inc., 2019. 11181–11191.
[11] Xiao YS, Xu RY, Wu LJ, Li JT, Qin T, Liu TY, Zhang M. AMOM: Adaptive masking over masking for conditional masked language
model. In: Proc. of the 37th AAAI Conf. on Artificial Intelligence. Washington: AAAI Press, 2023. 13789–13797. [doi: 10.1609/AAAI.
V37I11.26615]
[12] Duan SF, Zhao H, Zhang DD. Syntax-aware data augmentation for neural machine translation. IEEE/ACM Trans. on Audio, Speech, and
Language Processing, 2023, 31: 2988–2999. [doi: 10.1109/TASLP.2023.3301214]
[13] Mondal SK, Zhang HX, Kabir HMD, Ni K, Dai HN. Machine translation and its evaluation: A study. Artificial Intelligence Review,
2023, 56(9): 10137–10226. [doi: 10.1007/S10462-023-10423-5]

