Page 302 - 《软件学报》2026年第2期
P. 302

张建新 等: 依存句法信息增强的完全非自回归翻译                                                         781


                     [doi: 10.18653/V1/2022.NAACL-MAIN.129]
                 [59]  Kingma DP, Ba J. Adam: A method for stochastic optimization. arXiv:1412.6980, 2017.
                 [60]  Graves  A,  Fernández  S,  Gomez  F,  Schmidhuber  J.  Connectionist  temporal  classification:  Labelling  unsegmented  sequence  data  with
                     recurrent  neural  networks.  In:  Proc.  of  the  23rd  Int’l  Conf.  on  Machine  Learning.  Pittsburgh:  ACM,  2006.  369–376.  [doi:  10.1145/
                     1143844.1143891]
                 [61]  Shu R, Lee J, Nakayama H, Cho K. Latent-variable non-autoregressive neural machine translation with deterministic inference using a
                     delta posterior. In: Proc. of the 34th AAAI Conf. on Artificial Intelligence. New York: AAAI Press, 2020. 8846–8853. [doi: 10.1609/
                     AAAI.V34I05.6413]
                 [62]  Kasai J, Cross J, Ghazvininejad M, Gu JT. Parallel machine translation with disentangled context Transformer. arXiv:2001.05136, 2020.
                 [63]  Hao YC, He SL, Jiao WX, Tu ZP, Lyu M, Wang X. Multi-task learning with shared encoder for non-autoregressive machine translation.
                     In:  Proc.  of  the  2021  Conf.  of  the  North  American  Chapter  of  the  Association  for  Computational  Linguistics:  Human  Language
                     Technologies. ACL, 2021. 3989–3996. [doi: 10.18653/V1/2021.NAACL-MAIN.313]
                 [64]  Geng  XW,  Feng  XC,  Qin  B.  Learning  to  rewrite  for  non-autoregressive  neural  machine  translation.  In:  Proc.  of  the  2021  Conf.  on
                     Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 3297–3308. [doi: 10.18653/V1/2021.EMNLP-MAIN.265]
                 [65]  Huang XS, Pérez F, Volkovs M. Improving non-autoregressive translation models without distillation. In: Proc. of the 10th Int’l Conf. on
                     Learning Representations (ICLR 2022). 2022.
                 [66]  Chen XR, Duan SF, Liu GS. Improving non-autoregressive machine translation with error exposure and consistency regularization. In:
                     Proc. of the 13th National CCF Conf. on Natural Language Processing and Chinese Computing. Hangzhou: Springer, 2025. 240–252.
                     [doi: 10.1007/978-981-97-9437-9_19]
                 [67]  Li ZH, Lin Z, He D, Tian F, Qin T, Wang LW, Liu TY. Hint-based training for non-autoregressive machine translation. In: Proc. of the
                     2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural Language Processing (EMNLP-
                     IJCNLP). Hong Kong: ACL, 2019. 5708–5713. [doi: 10.18653/V1/D19-1573]
                 [68]  Sun ZQ, Li ZH, Wang HQ, He D, Lin Z, Deng ZH. Fast structured decoding for sequence models. In: Proc. of the 33rd Int’l Conf. on
                     Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2019. 3016–3026.
                 [69]  Ran Q, Lin YK, Li P, Zhou J. Guiding non-autoregressive neural machine translation decoding with reordering information. In: Proc. of
                     the 35th AAAI Conf. on Artificial Intelligence. Virtually: AAAI Press, 2021. 13727–13735. [doi: 10.1609/AAAI.V35I15.17618]
                 [70]  Song  J,  Kim  S,  Yoon  S.  AligNART:  Non-autoregressive  neural  machine  translation  by  jointly  learning  to  estimate  alignment  and
                     translate. In: Proc. of the 2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 1–14. [doi: 10.
                     18653/V1/2021.EMNLP-MAIN.1]
                 [71]  Du CX, Tu ZP, Jiang J. Order-agnostic cross entropy for non-autoregressive machine translation. arXiv:2106.05093, 2021.
                 [72]  Zhan JA, Chen Q, Chen BX, Wang W, Bai Y, Gao Y. DePA: Improving non-autoregressive translation with dependency-aware decoder.
                     In: Proc. of the 20th Int’l Conf. on Spoken Language Translation (IWSLT 2023). Toronto: ACL, 2023. 478–490. [doi: 10.18653/v1/2023.
                     iwslt-1.47]
                 [73]  Gui ST, Shao CZ, Ma ZR, Zhang XS, Chen YJ, Feng Y. Non-autoregressive machine translation with probabilistic context-free grammar.
                     In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2023. 5598–5615.
                 [74]  Huang F, Zhou H, Liu Y, Li H, Huang ML. Directed acyclic Transformer for non-autoregressive machine translation. In: Proc. of the 39th
                     Int’l Conf. on Machine Learning. Baltimore: PMLR, 2022. 9410–9428.
                 [75]  Saharia C, Chan W, Saxena S, Norouzi M. Non-autoregressive machine translation with latent alignments. In: Proc. of the 2020 Conf. on
                     Empirical Methods in Natural Language Processing (EMNLP). 2020. 1098–1108. [doi: 10.18653/V1/2020.EMNLP-MAIN.83]
                 [76]  Ran  Q,  Lin  YK,  Li  P,  Zhou  J.  Learning  to  recover  from  multi-modality  errors  for  non-autoregressive  neural  machine  translation.
                     arXiv:2006.05165, 2020.
                 [77]  Shannon CE. A mathematical theory of communication. The Bell System Technical Journal, 1948, 27(3): 379–423. [doi: 10.1002/J.1538-
                     7305.1948.TB01338.X]
                 [78]  Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional Transformers for language understanding. In: Proc.
                     of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
                     Minneapolis: ACL. 2019. 4171–4186. [doi: 10.18653/v1/N19-1423]
                 [79]  Levenshtein VI. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 1966, 10(8): 707–710.
                 [80]  Damerau FJ. A technique for computer detection and correction of spelling errors. Communications of the ACM, 1964, 7(3): 171–176.
                     [doi: 10.1145/363958.363994]
                 [81]  Salton G, Wong A, Yang CS. A vector space model for automatic indexing. Communications of the ACM, 1975, 18(11): 613–620. [doi:
                     10.1145/361219.361220]
   297   298   299   300   301   302   303   304   305   306   307