Page 36 - 《软件学报》2026年第3期
P. 36
宋子文 等: 向量数据库中近似最近邻搜索关键技术综述 999
augmented large language models. In: Proc. of the 30th ACM SIGKDD Conf. on Knowledge Discovery and Data Mining. Barcelona:
ACM, 2024. 6491–6501. [doi: 10.1145/3637528.3671470]
[8] Zhang HL, Ji XD, Chen YL, Fu FC, Miao XP, Nie XN, Chen WP, Cui B. PQCache: Product quantization-based KVCache for long
context LLM inference. Proc. of the ACM on Management of Data, 2025, 3(3): 201. [doi: 10.1145/3725338]
[9] Indyk P, Motwani R. Approximate nearest neighbors: Towards removing the of dimensionality. In: Proc. of the 30th Annual ACM
Symp. on Theory of Computing. Dallas: ACM, 1998. 604–613. [doi: 10.1145/276698.276876]
[10] Li W, Zhang Y, Sun YF, Wang W, Li MJ, Zhang WJ, Lin XM. Approximate nearest neighbor search on high dimensional data—
Experiments, analyses, and improvement. IEEE Trans. on Knowledge and Data Engineering, 2020, 32(8): 1475–1488. [doi: 10.1109/
TKDE.2019.2909204]
[11] Wang MZ, Xu XL, Yue Q, Wang YX. A comprehensive survey and experimental comparison of graph-based approximate nearest
neighbor search. Proc. of the VLDB Endowment, 2021, 14(11): 1964–1978. [doi: 10.14778/3476249.3476255]
[12] Azizi I, Echihabi K, Palpanas T. Graph-based vector search: An experimental evaluation of the state-of-the-art. Proc. of the ACM on
Management of Data, 2025, 3(1): 43. [doi: 10.1145/3709693]
[13] Wang JD, Zhang T, Song JK, Sebe N, Shen HT. A survey on learning to hash. IEEE Trans. on Pattern Analysis and Machine
Intelligence, 2018, 40(4): 769–790. [doi: 10.1109/TPAMI.2017.2699960]
[14] Cai D. A revisit of hashing algorithms for approximate nearest neighbor search. IEEE Trans. on Knowledge and Data Engineering,
2021, 33(6): 2337–2348. [doi: 10.1109/TKDE.2019.2953897]
[15] Matsui Y, Uchida Y, Jégou H, Satoh S. A survey of product quantization. ITE Trans. on Media Technology and Applications, 2018,
6(1): 2–10. [doi: 10.3169/mta.6.2]
[16] Aumüller M, Bernhardsson E, Faithfull A. ANN-benchmarks: A benchmarking tool for approximate nearest neighbor algorithms.
Information Systems, 2020, 87: 101374. [doi: 10.1016/j.is.2019.02.006]
[17] Fu C, Xiang C, Wang CX, Cai D. Fast approximate nearest neighbor search with the navigating spreading-out graph. Proc. of the VLDB
Endowment, 2019, 12(5): 461–474. [doi: 10.14778/3303753.3303754]
[18] Reimers N, Gurevych I. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In: Proc. of the 2019 Conf. on Empirical
Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural Language Processing (EMNLP-IJCNLP). Hong Kong:
ACL, 2019. 3982–3992. [doi: 10.18653/v1/D19-1410]
[19] OpenAI Platform. Vector embeddings. 2025. https://platform.openai.com/docs/guides/embeddings
[20] Pilehvar MT, Camacho-Collados J. Embeddings in Natural Language Processing: Theory and Advances in Vector Representations of
Meaning. Cham: Springer, 2021. 1–8. [doi: 10.1007/978-3-031-02177-0]
[21] Salton G, Buckley C. Term-weighting approaches in automatic text retrieval. Information Processing & Management, 1988, 24(5):
513–523. [doi: 10.1016/0306-4573(88)90021-0]
[22] Lassance C, Clinchant S. An efficiency study for SPLADE models. In: Proc. of the 45th Int’l ACM SIGIR Conf. on Research and
Development in Information Retrieval. Madrid: ACM, 2022. 2220–2226. [doi: 10.1145/3477495.3531833]
[23] Zhang XN, Liu XB, Song JK, Nie XS, Wang SH, Yin YL. Survey on hash learning for large-scale image retrieval. Ruan Jian Xue
Bao/Journal of Software, 2025, 36(1): 79–106 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7141.htm [doi: 10.
13328/j.cnki.jos.007141]
[24] Li TY, Yang XC, Ke YP, Wang B, Liu YN, Xu JX. Alleviating the inconsistency of multimodal data in cross-modal retrieval. In: Proc.
of the 40th IEEE Int’l Conf. on Data Engineering. Utrecht: IEEE, 2024. 4643–4656. [doi: 10.1109/ICDE60146.2024.00353]
[25] Robertson SE, Walker S, Jones S, Hancock-Beaulieu MM, Gatford M. Okapi at TREC-3. In: Proc. of the 3rd NIST Text Retrieval Conf.
(TREC3). Washington: NIST, 1996. 109–126.
[26] Bhattacharya A. Fundamentals of Database Indexing and Searching. New York: Chapman and Hall/CRC, 2014. [doi: 10.1201/b17767]
[27] Leskovec J, Rajaraman A, Ullman JD. Mining of Massive Datasets. 3rd ed., Cambridge: Cambridge University Press, 2020. 89–99.
[28] Malkov YA, Yashunin DA. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.
IEEE Trans. on Pattern Analysis and Machine Intelligence, 2020, 42(4): 824–836. [doi: 10.1109/TPAMI.2018.2889473]
[29] facebookresearch/faiss: A library for efficient similarity search and clustering of dense vectors. 2025. https://github.com/facebookresearch/
faiss
[30] Bentley JL. Multidimensional binary search trees used for associative searching. Communications of the ACM, 1975, 18(9): 509–517.
[doi: 10.1145/361002.361007]
[31] Huang Q, Feng JL, Zhang YK, Fang Q, Ng W. Query-aware locality-sensitive hashing for approximate nearest neighbor search. Proc. of
the VLDB Endowment, 2015, 9(1): 1–12. [doi: 10.14778/2850469.2850470]

