Page 36 - 《软件学报》2026年第3期
P. 36

宋子文 等: 向量数据库中近似最近邻搜索关键技术综述                                                       999


                      augmented large language models. In: Proc. of the 30th ACM SIGKDD Conf. on Knowledge Discovery and Data Mining. Barcelona:
                      ACM, 2024. 6491–6501. [doi: 10.1145/3637528.3671470]
                  [8]   Zhang HL, Ji XD, Chen YL, Fu FC, Miao XP, Nie XN, Chen WP, Cui B. PQCache: Product quantization-based KVCache for long
                      context LLM inference. Proc. of the ACM on Management of Data, 2025, 3(3): 201. [doi: 10.1145/3725338]
                  [9]   Indyk P, Motwani R. Approximate nearest neighbors: Towards removing the of dimensionality. In: Proc. of the 30th Annual ACM
                      Symp. on Theory of Computing. Dallas: ACM, 1998. 604–613. [doi: 10.1145/276698.276876]
                 [10]   Li W, Zhang Y, Sun YF, Wang W, Li MJ, Zhang WJ, Lin XM. Approximate nearest neighbor search on high dimensional data—
                      Experiments, analyses, and improvement. IEEE Trans. on Knowledge and Data Engineering, 2020, 32(8): 1475–1488. [doi: 10.1109/
                      TKDE.2019.2909204]
                 [11]   Wang  MZ,  Xu  XL,  Yue  Q,  Wang  YX.  A  comprehensive  survey  and  experimental  comparison  of  graph-based  approximate  nearest
                      neighbor search. Proc. of the VLDB Endowment, 2021, 14(11): 1964–1978. [doi: 10.14778/3476249.3476255]
                 [12]   Azizi I, Echihabi K, Palpanas T. Graph-based vector search: An experimental evaluation of the state-of-the-art. Proc. of the ACM on
                      Management of Data, 2025, 3(1): 43. [doi: 10.1145/3709693]
                 [13]   Wang  JD,  Zhang  T,  Song  JK,  Sebe  N,  Shen  HT.  A  survey  on  learning  to  hash.  IEEE  Trans.  on  Pattern  Analysis  and  Machine
                      Intelligence, 2018, 40(4): 769–790. [doi: 10.1109/TPAMI.2017.2699960]
                 [14]   Cai D. A revisit of hashing algorithms for approximate nearest neighbor search. IEEE Trans. on Knowledge and Data Engineering,
                      2021, 33(6): 2337–2348. [doi: 10.1109/TKDE.2019.2953897]
                 [15]   Matsui Y, Uchida Y, Jégou H, Satoh S. A survey of product quantization. ITE Trans. on Media Technology and Applications, 2018,
                      6(1): 2–10. [doi: 10.3169/mta.6.2]
                 [16]   Aumüller  M,  Bernhardsson  E,  Faithfull  A.  ANN-benchmarks:  A  benchmarking  tool  for  approximate  nearest  neighbor  algorithms.
                      Information Systems, 2020, 87: 101374. [doi: 10.1016/j.is.2019.02.006]
                 [17]   Fu C, Xiang C, Wang CX, Cai D. Fast approximate nearest neighbor search with the navigating spreading-out graph. Proc. of the VLDB
                      Endowment, 2019, 12(5): 461–474. [doi: 10.14778/3303753.3303754]
                 [18]   Reimers N, Gurevych I. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In: Proc. of the 2019 Conf. on Empirical
                      Methods in Natural Language Processing and the 9th Int’l Joint Conf. on Natural Language Processing (EMNLP-IJCNLP). Hong Kong:
                      ACL, 2019. 3982–3992. [doi: 10.18653/v1/D19-1410]
                 [19]   OpenAI Platform. Vector embeddings. 2025. https://platform.openai.com/docs/guides/embeddings
                 [20]   Pilehvar MT, Camacho-Collados J. Embeddings in Natural Language Processing: Theory and Advances in Vector Representations of
                      Meaning. Cham: Springer, 2021. 1–8. [doi: 10.1007/978-3-031-02177-0]
                 [21]   Salton  G,  Buckley  C.  Term-weighting  approaches  in  automatic  text  retrieval.  Information  Processing  &  Management,  1988,  24(5):
                      513–523. [doi: 10.1016/0306-4573(88)90021-0]
                 [22]   Lassance C, Clinchant S. An efficiency study for SPLADE models. In: Proc. of the 45th Int’l ACM SIGIR Conf. on Research and
                      Development in Information Retrieval. Madrid: ACM, 2022. 2220–2226. [doi: 10.1145/3477495.3531833]
                 [23]   Zhang XN, Liu XB, Song JK, Nie XS, Wang SH, Yin YL. Survey on hash learning for large-scale image retrieval. Ruan Jian Xue
                      Bao/Journal of Software, 2025, 36(1): 79–106 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7141.htm [doi: 10.
                      13328/j.cnki.jos.007141]
                 [24]   Li TY, Yang XC, Ke YP, Wang B, Liu YN, Xu JX. Alleviating the inconsistency of multimodal data in cross-modal retrieval. In: Proc.
                      of the 40th IEEE Int’l Conf. on Data Engineering. Utrecht: IEEE, 2024. 4643–4656. [doi: 10.1109/ICDE60146.2024.00353]
                 [25]   Robertson SE, Walker S, Jones S, Hancock-Beaulieu MM, Gatford M. Okapi at TREC-3. In: Proc. of the 3rd NIST Text Retrieval Conf.
                      (TREC3). Washington: NIST, 1996. 109–126.
                 [26]   Bhattacharya A. Fundamentals of Database Indexing and Searching. New York: Chapman and Hall/CRC, 2014. [doi: 10.1201/b17767]
                 [27]   Leskovec J, Rajaraman A, Ullman JD. Mining of Massive Datasets. 3rd ed., Cambridge: Cambridge University Press, 2020. 89–99.
                 [28]   Malkov YA, Yashunin DA. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.
                      IEEE Trans. on Pattern Analysis and Machine Intelligence, 2020, 42(4): 824–836. [doi: 10.1109/TPAMI.2018.2889473]
                 [29]   facebookresearch/faiss: A library for efficient similarity search and clustering of dense vectors. 2025. https://github.com/facebookresearch/
                      faiss
                 [30]   Bentley JL. Multidimensional binary search trees used for associative searching. Communications of the ACM, 1975, 18(9): 509–517.
                      [doi: 10.1145/361002.361007]
                 [31]   Huang Q, Feng JL, Zhang YK, Fang Q, Ng W. Query-aware locality-sensitive hashing for approximate nearest neighbor search. Proc. of
                      the VLDB Endowment, 2015, 9(1): 1–12. [doi: 10.14778/2850469.2850470]
   31   32   33   34   35   36   37   38   39   40   41