Page 94 - 《软件学报》2026年第3期
P. 94
李忠根 等: GPU 加速的高维向量聚类算法 1057
Symp. of Creative Computing (ISPAN-FCST-ISCC). Exeter: IEEE, 2017. 397–402. [doi: 10.1109/ISPAN-FCST-ISCC.2017.75]
[44] Gowanlock M, Rude CM, Blair DM, Li JD, Pankratius V. A hybrid approach for optimizing parallel clustering throughput using the
GPU. IEEE Trans. on Parallel and Distributed Systems, 2019, 30(4): 766–777. [doi: 10.1109/TPDS.2018.2869777]
[45] Loh WK, Yu H. Fast density-based clustering through dataset partition using graphics processing units. Information Sciences, 2015, 308:
94–112. [doi: 10.1016/j.ins.2014.10.023]
[46] Prokopenko A, Lebrun-Grandie D, Arndt D. Fast tree-based algorithms for DBSCAN for low-dimensional data on GPUs. In: Proc. of the
52nd Int’l Conf. on Parallel Processing. Salt Lake City: ACM, 2023. 503–512. [doi: 10.1145/3605573.3605594]
[47] Raschka S, Patterson J, Nolet C. Machine learning in Python: Main developments and technology trends in data science, machine
learning, and artificial intelligence. Information, 2020, 11(4): 193. [doi: 10.3390/info11040193]
[48] Dong W, Moses C, Li K. Efficient K-nearest neighbor graph construction for generic similarity measures. In: Proc. of the 20th Int’l Conf.
on World Wide Web. Hyderabad: ACM, 2011. 577–586. [doi: 10.1145/1963405.1963487]
[49] Lee JD, Batcher KE. Minimizing communication in the bitonic sort. IEEE Trans. on Parallel and Distributed Systems, 2000, 11(5):
459–474. [doi: 10.1109/71.852399]
[50] Yandex BA, Lempitsky V. Efficient indexing of billion-scale datasets of deep descriptors. In: Proc. of the 2016 IEEE Conf. on Computer
Vision and Pattern Recognition. Las Vegas: IEEE, 2016. 2055–2063. [doi: 10.1109/CVPR.2016.226]
[51] SIFT and GIST. 2023. http://corpus-texmex.irisa.fr
[52] Million Song dataset benchmarks. 2012. http://www.ifs.tuwien.ac.at/mir/msd/
[53] Zhu YF, Luo CY, Ma RY, Chen L, Mao YR, Gao YJ. Density-based data clustering algorithm in multi-metric spaces. Ruan Jian Xue
Bao/Journal of Software, 2025, 36(2): 851–873 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7177.htm [doi: 10.
13328/j.cnki.jos.007177]
[54] Liu JX, Zhang X. STK: Clustering method based on contrastive learning embedding. Computer Science, 2024, 51(11A): 240400011 (in
Chinese with English abstract). [doi: 10.11896/jsjkx.240400011]
[55] GOFMM. 2025. https://github.com/ChenhanYu/hmlp
[56] HNSWlib. 2025. https://github.com/nmslib/hnswlib
[57] KGraph. 2025. https://github.com/aaalgo/kgraph
[58] cuVS. 2025. https://github.com/rapidsai/cuvs
附中文参考文献
[11] 杜梦瑶, 李清明, 张淼, 陈曦, 李新梦, 尹全军, 纪守领. 面向隐私保护的用户评论基准数据集构建与大模型推理能力评估. 计算机学
报, 2025, 48(7): 1529–1550. [doi: 10.11897/SP.J.1016.2025.01529]
[53] 朱轶凡, 罗程阳, 马瑞遥, 陈璐, 毛玉仁, 高云君. 基于密度的多度量空间数据聚类算法. 软件学报, 2025, 36(2): 851–873. http://www.
jos.org.cn/1000-9825/7177.htm [doi: 10.13328/j.cnki.jos.007177]
[54] 刘晋霞, 张曦. STK: 基于对比学习嵌入的聚类方法. 计算机科学, 2024, 51(11A): 240400011. [doi: 10.11896/jsjkx.240400011]
作者简介
李忠根, 博士生, CCF 学生会员, 主要研究领域为新型硬件加速.
龚盛豪, 博士生, 主要研究领域为流式大数据管理与分析.
于浩然, 硕士生, 主要研究领域为向量数据库.
朱轶凡, 博士, 研究员, CCF 专业会员, 主要研究领域为大数据管理与分析.
柳晴, 博士, 研究员, 博士生导师, CCF 专业会员, 主要研究领域为数据库可用性, 图数据分析, 大数据.
高云君, 博士, 教授, 博士生导师, CCF 高级会员, 主要研究领域为数据库, 大数据管理与分析, DB 与 AI 融合.

