Page 346 - 《软件学报》2026年第7期
P. 346
胡晓雯 等: 基于 CUDA Core 和 Tensor Core 的 CTRU-Prime 高吞吐量实现 3031
[6] Schanck J, Whyte W, Zhang ZF. Circuit-extension handshakes for Tor achieving forward secrecy in a quantum world. Proc. on Privacy
Enhancing Technologies, 2016, 2016(4): 219–236. [doi: 10.1515/POPETS-2016-0037]
[7] OpenSSH. OpenSSH release notes—OpenSSH 9.0/9.0p1. 2022. https://www.openssh.com/releasenotes.html [doi: 10.1007/s10623-013-
9850-3]
[8] Augot D, Batina L, Bernstein DJ, Bos J, Buchmann J, Castryck W, Dunkelman O, GüneysuT, Gueron S, Hülsing A, Lange T, Mohamed
MSE, Rechberger C, Schwabe P, Sendrier N, Vercauteren F, Yang BY. Initial recommendations of long-term secure post-quantum
systems. 2015. https://pqcrypto.eu.org/docs/initial-recommendations.pdf
[9] Avanzi R, Bos J, Ducas L, Kiltz E, Lepoint T, Lyubashevsky V, Schanck JM, Schwabe P, Seiler G, Stehlé D. CRYSTALS-kyber
algorithm specifications and supporting documentation. NIST PQC Round, 2019, 2(4): 1–43.
[10] Weimerskirch A, Paar C. Generalizations of the Karatsuba algorithm for efficient implementations. IACR Cryptology ePrint Archive,
2006.224.
[11] Bernstein DJ, Chuengsatiansup C, Lange T, van Vredendaal C. NTRU prime: Reducing attack surface at low cost. In: Adams C,
Camenisch J, eds. Proc. of the 24th Int’l Conf. on Selected Areas in Cryptography (SAC 2017). Cham: Springer, 2018. 235–260. [doi: 10.
1007/978-3-319-72565-9_12]
[12] Liang ZC, Zhao XY, Fang BY, Zhao YL. Efficient and compact NTRU-based key encapsulation mechanism in large-galois-group prime-
degree prime-ideal number field. Ruan Jian Xue Bao/Journal of Software, 2025, 36(2): 747–775 (in Chinese with English abstract). http://
www.jos.org.cn/1000-9825/7161.htm [doi: 10.13328/j.cnki.jos.007161]
[13] Hafeez MA, Lee WK, Karmakar A, Hwang SO. TMVP-based polynomial convolution for Saber and Sable on GPU using CUDA-cores
and Tensor-cores. IACR Cryptology ePrint Archive, 2023.1541.
[14] Dai W, Sunar B, Schanck JM, Whyte W, Zhang Z. NTRU modular lattice signature scheme on CUDA GPUs. In: Proc. of the 2016 Int’l
Conf. on High Performance Computing & Simulation (HPCS 2016). Innsbruck: IEEE, 2016. 501–508. [doi: 10.1109/HPCSIM.2016.
7568376]
[15] Gupta N, Jati A, Chauhan AK, Chattopadhyay A. PQC acceleration using GPUs: FrodoKEM, NewHope, and Kyber. IEEE Trans. on
Parallel and Distributed Systems, 2021, 32(3): 575–586. [doi: 10.1109/TPDS.2020.3025691]
[16] Lee WK, Seo H, Zhang ZF, Hwang SO. TensorCrypto: High throughput acceleration of lattice-based cryptography using tensor core on
GPU. IEEE Access, 2022, 10: 20616–20632. [doi: 10.1109/ACCESS.2022.3152217]
[17] Sun SZ, Zhang R, Ma H. Efficient parallelism of post-quantum signature scheme SPHINCS. IEEE Trans. on Parallel and Distributed
Systems, 2020, 31(11): 2542–2555. [doi: 10.1109/TPDS.2020.2995562]
[18] Gao YW, Xu J, Wang HB. cuNH: Efficient GPU implementations of post-quantum KEM NewHope. IEEE Trans. on Parallel and
Distributed Systems, 2022, 33(3): 551–568. [doi: 10.1109/TPDS.2021.3097277]
[19] Shen SY, Yang H, Li WQ, Zhao YL. cuML-DSA: Optimized signing procedure and server-oriented GPU design for ML-DSA. IEEE
Trans. on Dependable and Secure Computing, 2025, 22(3): 2295–2307 [doi: 10.1109/TDSC.2024.3494835]
[20] Wan LP, Zheng FY, Fan G, Wei R, Gao LL, Wang YW, Lin JQ, Dong JK. A novel high-performance implementation of CRYSTALS-
Kyber with AI accelerator. In: Proc. of the 27th European Symp. on Research in Computer Security. Copenhagen: Springer, 2022.
514–534. [doi: 10.1007/978-3-031-17143-7_25]
[21] Zhou T, Zheng FY, Fan G, Wan LP, Tang WX, Song YX, Bian Y, Lin JQ. ConvKyber: Unleashing the power of AI accelerators for
faster Kyber with novel iteration-based approaches. IACR Trans. on Cryptographic Hardware and Embedded Systems, 2024, 2024(2):
25–63. [doi: 10.46586/TCHES.V2024.I2.25-63]
[22] Hafeez MA, Lee WK, Karmakar A, Hwang SO. High throughput acceleration of Scabbard key exchange and key encapsulation
mechanism using tensor core on GPU for IoT applications. IEEE Internet of Things Journal, 2023, 10(22): 19765–19781. [doi: 10.1109/
JIOT.2023.3282255]
[23] Kamal AA, Youssef AM. Enhanced implementation of the NTRUEncrypt algorithm using graphics cards. In: Proc. of the 1st Int’l Conf.
on Parallel, Distributed and Grid Computing (PDGC 2010). Solan: IEEE, 2010. 168–174. [doi: 10.1109/PDGC.2010.5679887]
[24] Hermans J, Vercauteren F, Preneel B. Speed records for NTRU. In: Proc. of the 10th Cryptographers’ Track at the RSA Conf. San
Francisco: Springer, 2010. 73–88. [doi: 10.1007/978-3-642-11925-5_6]
[25] Bai TY, Davis S, Li JJ, Gu Y, Jiang H. Accelerating NTRU encryption with graphics processing units. Int’l Journal of Networked and
Distributed Computing, 2014, 2(4): 250–258. [doi: 10.2991/IJNDC.2014.2.4.6]
[26] Akleylek S, Goi BM, Yap WS, Wong DCK, Lee WK. Fast NTRU encryption in GPU for secure IOP communication in post-quantum era.
In: Proc. of the 2018 IEEE SmartWorld, Ubiquitous Intelligence & Computing, Advanced & Trusted Computing, Scalable Computing &
Communications, Cloud & Big Data Computing, Internet of People and Smart City Innovation (SmartWorld/SCALCOM/UIC/ATC/

