Page 72 - 《软件学报》2026年第3期
P. 72
周依杰 等: GoVector: I/O-高效的高维向量近邻查询缓存策略 1035
[2] Xu YM, Liang HY, Li J, Xu ST, Chen Q, Zhang QX, Li C, Yang ZY, Yang F, Yang YQ, Cheng P, Yang M. SPFresh: Incremental
in-place update for billion-scale vector search. In: Proc. of the 29th Symp. on Operating Systems Principles. Koblenz: ACM, 2023.
545–561. [doi: 10.1145/3600006.3613166]
[3] Grbovic M, Cheng HB. Real-time personalization using embeddings for search ranking at airbnb. In: Proc. of the 24th ACM SIGKDD Int’l
Conf. on Knowledge Discovery & Data Mining. London: ACM, 2018. 311–320. [doi: 10.1145/3219819.3219885]
[4] Huang JT, Sharma A, Sun SY, Xia L, Zhang D, Pronin P, Padmanabhan J, Ottaviano G, Yang LJ. Embedding-based retrieval in facebook
search. In: Proc. of the 26th ACM SIGKDD Int’l Conf. on Knowledge Discovery & Data Mining. New York: ACM, 2020. 2553–2561.
[doi: 10.1145/3394486.3403305]
[5] Okura S, Tagami Y, Ono S, Tajima A. Embedding-based news recommendation for millions of users. In: Proc. of the 23rd ACM
SIGKDD Int’l Conf. on Knowledge Discovery and Data Mining. Halifax: ACM, 2017. 1933–1942. [doi: 10.1145/3097983.3098108]
[6] Covington P, Adams J, Sargin E. Deep neural networks for YouTube recommendations. In: Proc. of the 10th ACM Conf. on
Recommender Systems. Boston: ACM, 2016. 191–198. [doi: 10.1145/2959100.2959190]
[7] Lyu Y, Li ZY, Niu SM, Xiong FY, Tang B, Wang WJ, Wu H, Liu HY, Xu T, Chen EH. CRUD-RAG: A comprehensive Chinese
benchmark for retrieval-augmented generation of large language models. ACM Trans. on Information Systems, 2025, 43(2): 41. [doi: 10.
1145/3701228]
[8] Gao JY, Long C. High-dimensional approximate nearest neighbor search: With reliable and efficient distance comparison operations.
Proc. of the ACM on Management of Data, 2023, 1(2): 137. [doi: 10.1145/3589282]
[9] Malkov YA, Yashunin DA. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.
IEEE Trans. on Pattern Analysis and Machine Intelligence, 2020, 42(4): 824–836. [doi: 10.1109/TPAMI.2018.2889473]
[10] Fu C, Xiang C, Wang CX, Cai D. Fast approximate nearest neighbor search with the navigating spreading-out graph. Proc. of the VLDB
Endowment, 2019, 12(5): 461–474. [doi: 10.14778/3303753.3303754]
[11] Subramanya SJ, Devvrit, Kadekodi R, Krishaswamy R, Simhadri HV. DiskANN: Fast accurate billion-point nearest neighbor search on a
single node. In: Proc. of the 33rd Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2019. 1233.
[doi: 10.5555/3454287.3455520]
[12] Chen Q, Zhao B, Wang HD, Li MQ, Liu CJ, Li ZZ, Yang M, Wang JD. SPANN: Highly-efficient billion-scale approximate nearest
neighbor search. In: Proc. of the 35th Int’l Conf. on Neural Information Processing Systems. Curran Associates Inc., 2021. 398. [doi: 10.
5555/3540261.3540659]
[13] Sun YF, Wang W, Qin JB, Zhang Y, Lin XM. SRS: Solving c-approximate nearest neighbor queries in high dimensional euclidean space
with a tiny index. Proc. of the VLDB Endowment, 2014, 8(1): 1–12. [doi: 10.14778/2735461.2735462]
[14] Bentley JL. Multidimensional binary search trees used for associative searching. Communications of the ACM, 1975, 18(9): 509–517.
[doi: 10.1145/361002.361007]
[15] Arora A, Sinha S, Kumar P, Bhattacharya A. HD-index: Pushing the scalability-accuracy boundary for approximate kNN search in high-
dimensional spaces. Proc. of the VLDB Endowment, 2018, 11(8): 906–919. [doi: 10.14778/3204028.3204034]
[16] Wang MZ, Xu WZ, Yi XM, Wu SL, Peng ZY, Ke XY, Gao YJ, Xu XL, Guo RT, Xie C. Starling: An I/O-efficient disk-resident graph
index framework for high-dimensional vector similarity search on data segment. Proc. of the ACM on Management of Data, 2024, 2(1):
14. [doi: 10.1145/3639269]
[17] MacAvaney S, Mallia A, Tonellotto N. Efficient constant-space multi-vector retrieval. In: Proc. of the 47th European Conf. on
Information Retrieval. Lucca: Springer, 2025. 237–245. [doi: 10.1007/978-3-031-88714-7_22]
[18] Wang JG, Yi XM, Guo RT, Jin H, Xu P, Li SJ, Wang XY, Guo XZ, Li CM, Xu XH, Yu K, Yuan YX, Zou YH, Long JQ, Cai YD, Li ZX,
Zhang ZF, Mo YH, Gu J, Jiang RY, Wei Y, Xie C. Milvus: A purpose-built vector data management system. In: Proc. of the 2021 Int’l
Conf. on Management of Data. New York: ACM, 2021. 2614–2627. [doi: 10.1145/3448016.3457550]
[19] Fu C, Wang CX, Cai D. High dimensional similarity search with satellite system graph: Efficiency, scalability, and unindexed query
compatibility. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2022, 44(8): 4139–4150. [doi: 10.1109/TPAMI.2021.3067706]
[20] Ge TZ, He KM, Ke QF, Sun J. Optimized product quantization. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2014, 36(4):
744–755. [doi: 10.1109/TPAMI.2013.240]
[21] Yin XZ, Gao C, Zhao ZJ, Gupta R. PANNS: Enhancing graph-based approximate nearest neighbor search through recency-aware
construction and parameterized search. In: Proc. of the 30th ACM SIGPLAN Annual Symp. on Principles and Practice of Parallel
Programming. Las Vegas: ACM, 2025. 369–381. [doi: 10.1145/3710848.3710867]
[22] Ram P, Sinha K. Revisiting kd-tree for nearest neighbor search. In: Proc. of the 25th ACM SIGKDD Int’l Conf. on Knowledge Discovery
& Data Mining. Anchorage: ACM, 2019. 1378–1388. [doi: 10.1145/3292500.3330875]

