Page 139 - 《软件学报》2026年第3期
P. 139
1102 软件学报 2026 年第 37 卷第 3 期
2025, 36(4): 1665–1691 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7245.htm [doi: 10.13328/j.cnki.jos.007245]
[8] Almeida F, Xexéo G. Word embeddings: A survey. arXiv:1901.09069, 2019.
[9] Grohe M. word2vec, node2vec, graph2vec, X2vec: Towards a theory of vector embeddings of structured data. In: Proc. of the 39th ACM
SIGMOD-SIGACT-SIGAI Symp. on Principles of Database Systems. Portland: ACM, 2020. 1–16. [doi: 10.1145/3375395.3387641]
[10] Pan JJ, Wang JG, Li GL. Survey of vector database management systems. The VLDB Journal, 2024, 33(5): 1591–1615. [doi: 10.1007/
s00778-024-00864-x]
[11] Pan JJ, Wang JG, Li GL. Vector database management techniques and systems. In: Proc. of the Companion of the 2024 Int’l Conf. on
Management of Data. Santiago: ACM, 2024. 597–604. [doi: 10.1145/3626246.3654691]
[12] Zhao WX, Zhou K, Li JY, Tang TY, Wang XL, Hou YP, Min YQ, Zhang BC, Zhang JJ, Dong ZC, Du YF, Yang C, Chen YS, Chen ZP,
Jiang JH, Ren RY, Li YF, Tang XY, Liu ZK, Liu PY, Nie JY, Wen JR. A survey of large language models. arXiv:2303.18223, 2023.
[13] Zhu XP, Yao HD, Liu J, Xiong XK. Review of evolution of large language model algorithms. ZTE Technology Journal, 2024, 30(2):
9–20 (in Chinese with English abstract). [doi: 10.12142/ZTETJ.202402003]
[14] Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, Küttler H, Lewis M, Yih WT, Rocktäschel T, Riedel S, Kiela D. Retrieval-
augmented generation for knowledge-intensive NLP tasks. In: Proc. of the 34th Int’l Conf. on Neural Information Processing Systems.
Vancouver: Curran Associates Inc., 2020. 793.
[15] Gao YF, Xiong Y, Gao XY, Jia KX, Pan JL, Bi YX, Dai Y, Sun JW, Wang M, Wang HF. Retrieval-augmented generation for large
language models: A survey. arXiv:2312.10997, 2023.
[16] Fan WQ, Ding YJ, Ning LB, Wang SJ, Li HY, Yin DW, Chua TS, Li Q. A survey on RAG meeting LLMS: Towards retrieval-augmented
large language models. In: Proc. of the 30th ACM SIGKDD Conf. on Knowledge Discovery and Data Mining. Barcelona: ACM, 2024.
6491–6501. [doi: 10.1145/3637528.3671470]
[17] Ren RY, Wang YH, Qu YQ, Zhao WX, Liu J, Wu H, Wen JR, Wang HF. Investigating the factual knowledge boundary of large language
models with retrieval augmentation. In: Proc. of the 31st Int’l Conf. on Computational Linguistics. Abu Dhabi: ACL, 2025. 3697–3715.
[18] Liu ZY, Wang PJ, Song XB, Zhang X, Jiang BB. Survey on hallucinations in large language models. Ruan Jian Xue Bao/Journal of
Software, 2025, 36(3): 1152–1185 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/7242.htm [doi: 10.13328/j.cnki.jos.
007242]
[19] Bentley JL. Multidimensional binary search trees used for associative searching. Communications of the ACM, 1975, 18(9): 509–517.
[doi: 10.1145/361002.361007]
[20] Dolatshah M, Hadian A, Minaei-Bidgoli B. Ball*-tree: Efficient spatial indexing for constrained nearest-neighbor search in metric spaces.
arXiv:1511.00628, 2015.
[21] Malkov Y, Ponomarenko A, Logvinov A, Krylov V. Approximate nearest neighbor algorithm based on navigable small world graphs.
Information Systems, 2014, 45: 61–68. [doi: 10.1016/j.is.2013.10.006]
[22] Wang MZ, Xu XL, Yue Q, Wang YX. A comprehensive survey and experimental comparison of graph-based approximate nearest
neighbor search. Proc. of the VLDB Endowment, 2021, 14(11): 1964–1978. [doi: 10.14778/3476249.3476255]
[23] Wang JD, Li SP. Query-driven iterated neighborhood graph search for large scale indexing. In: Proc. of the 20th ACM Int’l Conf. on
Multimedia. Nara: ACM, 2012. 179–188. [doi: 10.1145/2393347.2393378]
[24] Gionis A, Indyk P, Motwani R. Similarity search in high dimensions via hashing. In: Proc. of the 25th Int’l Conf. on Very Large Data
Bases. Edinburgh: Morgan Kaufmann Publishers Inc., 1999. 518–529.
[25] Datar M, Immorlica N, Indyk P, Mirrokni VS. Locality-sensitive hashing scheme based on p-stable distributions. In: Proc. of the 12th
Annual Symp. on Computational Geometry. New York: ACM, 2004. 253–262. [doi: 10.1145/997817.997857]
[26] Blumer A, Blumer J, Haussler D, McConnell R, Ehrenfeucht A. Complete inverted files for efficient text retrieval and analysis. Journal of
the ACM (JACM), 1987, 34(3): 578–595. [doi: 10.1145/28869.28873]
[27] Gray R. Vector quantization. IEEE ASSP Magazine, 1984, 1(2): 4–29. [doi: 10.1109/MASSP.1984.1162229]
[28] Pan ZB, Wang LZ, Wang Y, Liu YC. Product quantization with dual codebooks for approximate nearest neighbor search. Neurocomputing,
2020, 401: 59–68. [doi: 10.1016/j.neucom.2020.03.016]
[29] Malkov YA, Yashunin DA. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.
IEEE Trans. on Pattern Analysis and Machine Intelligence, 2020, 42(4): 824–836. [doi: 10.1109/TPAMI.2018.2889473]
[30] Chen Q, Zhao B, Wang HD, Li MQ, Liu CJ, Li ZZ, Yang M, Wang JD. SPANN: Highly-efficient billion-scale approximate nearest
neighbor search. In: Proc. of the 35th Int’l Conf. on Neural Information Processing Systems. Curran Associates Inc., 2021. 398.
[31] Echihabi K, Zoumpatianos K, Palpanas T. New trends in high-D vector similarity search: AI-driven, progressive, and distributed. Proc. of
the VLDB Endowment, 2021, 14(12): 3198–3201. [doi: 10.14778/3476311.3476407]

