Page 299 - 《软件学报》2026年第2期
P. 299

778                                                        软件学报  2026  年第  37  卷第  2  期


                    Grammar Error Count: 该指标使用  LanguageTool 工具对句子进行错误检测, 统计其中的错误数量, 主要包括
                 主谓不一致、句子结构不完整、拼写错误等. 该指标以错误数量计, 数值越低表示句子质量越高                              [84] .
                    上述指标从结构合法性、流畅性判断和细节检测这                   3  个层面客观地验证我们方法在提升目标端句法准确性
                 方面的有效性, 且在业界评估实践中均为常用手段. 实验结果如表                    11  所示. 从表中最后两行所有数据测试的平均
                 值可见, 在  3  项指标上, SynNAT  生成的译文整体优于        GLAT, 表明源端句法建模在目标句法结构的生成中起到了
                 积极作用.

                                          表 11 蒸馏的    IWSLT14 DE→EN  测试样本评估

                     模型               句子               Parse Validity  Acceptability Score  Grammar Error Count
                    GLAT        the only country country.  1             0.760 4              3
                   SynNAT      the only country in the world.  1         0.948 2              2
                    GLAT         she did for three years.  1             0.859 9              2
                   SynNAT       she did it for three years.  1           0.981 7              2
                    GLAT            all data avg         0.482 4         0.480 4            7.469 0
                   SynNAT           all data avg         0.493 4         0.492 6            7.378 7

                  5   结 论

                    本研究重新探讨了依存句法信息增强源端是否能够提升                    Fully NAT  的性能. 通过将源句的丰富依存句法结构
                 显式编码到    Fully NAT  模型中, 我们提出的    SynNAT  模型显著提高了翻译质量, 同时不会产生明显的解码速度损
                 失. 实验结果表明, 该方法在多个翻译基准测试中表现出色, 显著提升了                     Fully NAT  模型的翻译性能, 验证了利用
                 GCN  对依存句法信息进行编码的有效性和潜力. 该创新方法不仅为                    NAT  模型的发展提供了宝贵的新见解, 也为
                 实现高效且高质量的非自回归翻译奠定了基础.


                 References:
                  [1]  Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I. Attention is all you need. In: Proc. of the
                     31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
                  [2]  OpenAI. GPT-4 technical report. arXiv:2303.08774, 2024.
                  [3]  Dubey A, Jauhri A, Pandey A, et al. The llama 3 herd of models. arXiv:2407.21783, 2024.
                  [4]  Lee J, Stevens N, Han SC, Song M. A survey of large language models in finance (FinLLMs). arXiv:2402.02315, 2024.
                  [5]  Cai TL, Li YH, Geng ZY, Peng HW, Lee JD, Chen DM, Dao T. Medusa: Simple LLM inference acceleration framework with multiple
                     decoding heads. arXiv:2401.10774, 2024.
                  [6]  Gu JT, Bradbury J, Xiong CM, Li VOK, Socher R. Non-autoregressive neural machine translation. arXiv:1711.02281, 2018.
                  [7]  Guo  P,  Xiao  YS,  Li  JT,  Ji  YX,  Zhang  M.  Isotropy-enhanced  conditional  masked  language  models.  In:  Proc.  of  the  2023  Conf.  on
                     Empirical Methods in Natural Language Processing. Singapore: ACL, 2023. 8278–8289. [doi: 10.18653/v1/2023.findings-emnlp.555]
                  [8]  Xiao YS, Wu LJ, Guo JL, Li JT, Zhang M, Qin T, Liu TY. A survey on non-autoregressive generation for neural machine translation and
                     beyond. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2023, 45(10): 11407–11427. [doi: 10.1109/TPAMI.2023.3277122]
                  [9]  Ghazvininejad M, Levy O, Liu YH, Zettlemoyer L. Mask-predict: Parallel decoding of conditional masked language models. In: Proc. of
                     the 2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural Language Processing
                     (EMNLP-IJCNLP). Hong Kong: ACL, 2019. 6112–6121. [doi: 10.18653/V1/D19-1633]
                 [10]  Gu  JT,  Wang  CH,  Zhao  JB.  Levenshtein  Transformer.  In:  Proc.  of  the  33rd  Int’l  Conf.  on  Neural  Information  Processing  Systems.
                     Vancouver: Curran Associates Inc., 2019. 11181–11191.
                 [11]  Xiao YS, Xu RY, Wu LJ, Li JT, Qin T, Liu TY, Zhang M. AMOM: Adaptive masking over masking for conditional masked language
                     model. In: Proc. of the 37th AAAI Conf. on Artificial Intelligence. Washington: AAAI Press, 2023. 13789–13797. [doi: 10.1609/AAAI.
                     V37I11.26615]
                 [12]  Duan SF, Zhao H, Zhang DD. Syntax-aware data augmentation for neural machine translation. IEEE/ACM Trans. on Audio, Speech, and
                     Language Processing, 2023, 31: 2988–2999. [doi: 10.1109/TASLP.2023.3301214]
                 [13]  Mondal SK, Zhang HX, Kabir HMD, Ni K, Dai HN. Machine translation and its evaluation: A study. Artificial Intelligence Review,
                     2023, 56(9): 10137–10226. [doi: 10.1007/S10462-023-10423-5]
   294   295   296   297   298   299   300   301   302   303   304