Page 302 - 《软件学报》2026年第2期
P. 302
张建新 等: 依存句法信息增强的完全非自回归翻译 781
[doi: 10.18653/V1/2022.NAACL-MAIN.129]
[59] Kingma DP, Ba J. Adam: A method for stochastic optimization. arXiv:1412.6980, 2017.
[60] Graves A, Fernández S, Gomez F, Schmidhuber J. Connectionist temporal classification: Labelling unsegmented sequence data with
recurrent neural networks. In: Proc. of the 23rd Int’l Conf. on Machine Learning. Pittsburgh: ACM, 2006. 369–376. [doi: 10.1145/
1143844.1143891]
[61] Shu R, Lee J, Nakayama H, Cho K. Latent-variable non-autoregressive neural machine translation with deterministic inference using a
delta posterior. In: Proc. of the 34th AAAI Conf. on Artificial Intelligence. New York: AAAI Press, 2020. 8846–8853. [doi: 10.1609/
AAAI.V34I05.6413]
[62] Kasai J, Cross J, Ghazvininejad M, Gu JT. Parallel machine translation with disentangled context Transformer. arXiv:2001.05136, 2020.
[63] Hao YC, He SL, Jiao WX, Tu ZP, Lyu M, Wang X. Multi-task learning with shared encoder for non-autoregressive machine translation.
In: Proc. of the 2021 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language
Technologies. ACL, 2021. 3989–3996. [doi: 10.18653/V1/2021.NAACL-MAIN.313]
[64] Geng XW, Feng XC, Qin B. Learning to rewrite for non-autoregressive neural machine translation. In: Proc. of the 2021 Conf. on
Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 3297–3308. [doi: 10.18653/V1/2021.EMNLP-MAIN.265]
[65] Huang XS, Pérez F, Volkovs M. Improving non-autoregressive translation models without distillation. In: Proc. of the 10th Int’l Conf. on
Learning Representations (ICLR 2022). 2022.
[66] Chen XR, Duan SF, Liu GS. Improving non-autoregressive machine translation with error exposure and consistency regularization. In:
Proc. of the 13th National CCF Conf. on Natural Language Processing and Chinese Computing. Hangzhou: Springer, 2025. 240–252.
[doi: 10.1007/978-981-97-9437-9_19]
[67] Li ZH, Lin Z, He D, Tian F, Qin T, Wang LW, Liu TY. Hint-based training for non-autoregressive machine translation. In: Proc. of the
2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural Language Processing (EMNLP-
IJCNLP). Hong Kong: ACL, 2019. 5708–5713. [doi: 10.18653/V1/D19-1573]
[68] Sun ZQ, Li ZH, Wang HQ, He D, Lin Z, Deng ZH. Fast structured decoding for sequence models. In: Proc. of the 33rd Int’l Conf. on
Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2019. 3016–3026.
[69] Ran Q, Lin YK, Li P, Zhou J. Guiding non-autoregressive neural machine translation decoding with reordering information. In: Proc. of
the 35th AAAI Conf. on Artificial Intelligence. Virtually: AAAI Press, 2021. 13727–13735. [doi: 10.1609/AAAI.V35I15.17618]
[70] Song J, Kim S, Yoon S. AligNART: Non-autoregressive neural machine translation by jointly learning to estimate alignment and
translate. In: Proc. of the 2021 Conf. on Empirical Methods in Natural Language Processing. Punta Cana: ACL, 2021. 1–14. [doi: 10.
18653/V1/2021.EMNLP-MAIN.1]
[71] Du CX, Tu ZP, Jiang J. Order-agnostic cross entropy for non-autoregressive machine translation. arXiv:2106.05093, 2021.
[72] Zhan JA, Chen Q, Chen BX, Wang W, Bai Y, Gao Y. DePA: Improving non-autoregressive translation with dependency-aware decoder.
In: Proc. of the 20th Int’l Conf. on Spoken Language Translation (IWSLT 2023). Toronto: ACL, 2023. 478–490. [doi: 10.18653/v1/2023.
iwslt-1.47]
[73] Gui ST, Shao CZ, Ma ZR, Zhang XS, Chen YJ, Feng Y. Non-autoregressive machine translation with probabilistic context-free grammar.
In: Proc. of the 37th Int’l Conf. on Neural Information Processing Systems. New Orleans: Curran Associates Inc., 2023. 5598–5615.
[74] Huang F, Zhou H, Liu Y, Li H, Huang ML. Directed acyclic Transformer for non-autoregressive machine translation. In: Proc. of the 39th
Int’l Conf. on Machine Learning. Baltimore: PMLR, 2022. 9410–9428.
[75] Saharia C, Chan W, Saxena S, Norouzi M. Non-autoregressive machine translation with latent alignments. In: Proc. of the 2020 Conf. on
Empirical Methods in Natural Language Processing (EMNLP). 2020. 1098–1108. [doi: 10.18653/V1/2020.EMNLP-MAIN.83]
[76] Ran Q, Lin YK, Li P, Zhou J. Learning to recover from multi-modality errors for non-autoregressive neural machine translation.
arXiv:2006.05165, 2020.
[77] Shannon CE. A mathematical theory of communication. The Bell System Technical Journal, 1948, 27(3): 379–423. [doi: 10.1002/J.1538-
7305.1948.TB01338.X]
[78] Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional Transformers for language understanding. In: Proc.
of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
Minneapolis: ACL. 2019. 4171–4186. [doi: 10.18653/v1/N19-1423]
[79] Levenshtein VI. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 1966, 10(8): 707–710.
[80] Damerau FJ. A technique for computer detection and correction of spelling errors. Communications of the ACM, 1964, 7(3): 171–176.
[doi: 10.1145/363958.363994]
[81] Salton G, Wong A, Yang CS. A vector space model for automatic indexing. Communications of the ACM, 1975, 18(11): 613–620. [doi:
10.1145/361219.361220]

