Page 72 - 《软件学报》2026年第3期
P. 72

周依杰 等: GoVector: I/O-高效的高维向量近邻查询缓存策略                                            1035


                  [2]   Xu YM, Liang HY, Li J, Xu ST, Chen Q, Zhang QX, Li C, Yang ZY, Yang F, Yang YQ, Cheng P, Yang M. SPFresh: Incremental
                     in-place  update  for  billion-scale  vector  search.  In:  Proc.  of  the  29th  Symp.  on  Operating  Systems  Principles.  Koblenz:  ACM,  2023.
                     545–561. [doi: 10.1145/3600006.3613166]
                  [3]   Grbovic M, Cheng HB. Real-time personalization using embeddings for search ranking at airbnb. In: Proc. of the 24th ACM SIGKDD Int’l
                     Conf. on Knowledge Discovery & Data Mining. London: ACM, 2018. 311–320. [doi: 10.1145/3219819.3219885]
                  [4]   Huang JT, Sharma A, Sun SY, Xia L, Zhang D, Pronin P, Padmanabhan J, Ottaviano G, Yang LJ. Embedding-based retrieval in facebook
                     search. In: Proc. of the 26th ACM SIGKDD Int’l Conf. on Knowledge Discovery & Data Mining. New York: ACM, 2020. 2553–2561.
                     [doi: 10.1145/3394486.3403305]
                  [5]   Okura  S,  Tagami  Y,  Ono  S,  Tajima  A.  Embedding-based  news  recommendation  for  millions  of  users.  In:  Proc.  of  the  23rd  ACM
                     SIGKDD Int’l Conf. on Knowledge Discovery and Data Mining. Halifax: ACM, 2017. 1933–1942. [doi: 10.1145/3097983.3098108]
                  [6]   Covington  P,  Adams  J,  Sargin  E.  Deep  neural  networks  for  YouTube  recommendations.  In:  Proc.  of  the  10th  ACM  Conf.  on
                     Recommender Systems. Boston: ACM, 2016. 191–198. [doi: 10.1145/2959100.2959190]
                  [7]   Lyu  Y,  Li  ZY,  Niu  SM,  Xiong  FY,  Tang  B,  Wang  WJ,  Wu  H,  Liu  HY,  Xu  T,  Chen  EH.  CRUD-RAG:  A  comprehensive  Chinese
                     benchmark for retrieval-augmented generation of large language models. ACM Trans. on Information Systems, 2025, 43(2): 41. [doi: 10.
                     1145/3701228]
                  [8]   Gao JY, Long C. High-dimensional approximate nearest neighbor search: With reliable and efficient distance comparison operations.
                     Proc. of the ACM on Management of Data, 2023, 1(2): 137. [doi: 10.1145/3589282]
                  [9]   Malkov YA, Yashunin DA. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.
                     IEEE Trans. on Pattern Analysis and Machine Intelligence, 2020, 42(4): 824–836. [doi: 10.1109/TPAMI.2018.2889473]
                 [10]   Fu C, Xiang C, Wang CX, Cai D. Fast approximate nearest neighbor search with the navigating spreading-out graph. Proc. of the VLDB
                     Endowment, 2019, 12(5): 461–474. [doi: 10.14778/3303753.3303754]
                 [11]   Subramanya SJ, Devvrit, Kadekodi R, Krishaswamy R, Simhadri HV. DiskANN: Fast accurate billion-point nearest neighbor search on a
                     single node. In: Proc. of the 33rd Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2019. 1233.
                     [doi: 10.5555/3454287.3455520]
                 [12]   Chen Q, Zhao B, Wang HD, Li MQ, Liu CJ, Li ZZ, Yang M, Wang JD. SPANN: Highly-efficient billion-scale approximate nearest
                     neighbor search. In: Proc. of the 35th Int’l Conf. on Neural Information Processing Systems. Curran Associates Inc., 2021. 398. [doi: 10.
                     5555/3540261.3540659]
                 [13]   Sun YF, Wang W, Qin JB, Zhang Y, Lin XM. SRS: Solving c-approximate nearest neighbor queries in high dimensional euclidean space
                     with a tiny index. Proc. of the VLDB Endowment, 2014, 8(1): 1–12. [doi: 10.14778/2735461.2735462]
                 [14]   Bentley JL. Multidimensional binary search trees used for associative searching. Communications of the ACM, 1975, 18(9): 509–517.
                     [doi: 10.1145/361002.361007]
                 [15]   Arora A, Sinha S, Kumar P, Bhattacharya A. HD-index: Pushing the scalability-accuracy boundary for approximate kNN search in high-
                     dimensional spaces. Proc. of the VLDB Endowment, 2018, 11(8): 906–919. [doi: 10.14778/3204028.3204034]
                 [16]   Wang MZ, Xu WZ, Yi XM, Wu SL, Peng ZY, Ke XY, Gao YJ, Xu XL, Guo RT, Xie C. Starling: An I/O-efficient disk-resident graph
                     index framework for high-dimensional vector similarity search on data segment. Proc. of the ACM on Management of Data, 2024, 2(1):
                     14. [doi: 10.1145/3639269]
                 [17]   MacAvaney  S,  Mallia  A,  Tonellotto  N.  Efficient  constant-space  multi-vector  retrieval.  In:  Proc.  of  the  47th  European  Conf.  on
                     Information Retrieval. Lucca: Springer, 2025. 237–245. [doi: 10.1007/978-3-031-88714-7_22]
                 [18]   Wang JG, Yi XM, Guo RT, Jin H, Xu P, Li SJ, Wang XY, Guo XZ, Li CM, Xu XH, Yu K, Yuan YX, Zou YH, Long JQ, Cai YD, Li ZX,
                     Zhang ZF, Mo YH, Gu J, Jiang RY, Wei Y, Xie C. Milvus: A purpose-built vector data management system. In: Proc. of the 2021 Int’l
                     Conf. on Management of Data. New York: ACM, 2021. 2614–2627. [doi: 10.1145/3448016.3457550]
                 [19]   Fu C, Wang CX, Cai D. High dimensional similarity search with satellite system graph: Efficiency, scalability, and unindexed query
                     compatibility. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2022, 44(8): 4139–4150. [doi: 10.1109/TPAMI.2021.3067706]
                 [20]   Ge TZ, He KM, Ke QF, Sun J. Optimized product quantization. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2014, 36(4):
                     744–755. [doi: 10.1109/TPAMI.2013.240]
                 [21]   Yin  XZ,  Gao  C,  Zhao  ZJ,  Gupta  R.  PANNS:  Enhancing  graph-based  approximate  nearest  neighbor  search  through  recency-aware
                     construction  and  parameterized  search.  In:  Proc.  of  the  30th  ACM  SIGPLAN  Annual  Symp.  on  Principles  and  Practice  of  Parallel
                     Programming. Las Vegas: ACM, 2025. 369–381. [doi: 10.1145/3710848.3710867]
                 [22]   Ram P, Sinha K. Revisiting kd-tree for nearest neighbor search. In: Proc. of the 25th ACM SIGKDD Int’l Conf. on Knowledge Discovery
                     & Data Mining. Anchorage: ACM, 2019. 1378–1388. [doi: 10.1145/3292500.3330875]
   67   68   69   70   71   72   73   74   75   76   77