Page 155 - 《软件学报》2026年第3期
P. 155

1118                                                       软件学报  2026  年第  37  卷第  3  期


                    如图  7  中实验结果所示, 随着批量插入规模的提升, WRVQ             方法的索引增量操作在时间和空间开销上均表现
                 出良好的线性可扩展性, 无明显的性能突变或瓶颈, 验证了                 WRVQ  对增量索引构建的良好适应性.

                                                                                                2.00
                      0.084
                               批量时间变化量                                                          1.75
                     平均单个批量时间变化 (s)  0.072                                                      1.25 平均单个批量 CUDA 占用变化 (MB)
                      0.080
                               CUDA 内存变化量
                                                                                                1.50
                      0.076
                                                                                                1.00
                      0.068
                                                                                                0.75
                      0.064
                      0.060
                                                                                                0.25
                      0.056                                                                     0.50
                                   5 000        10 000       15 000      20 000       25 000
                                                         批量样例数
                                        图 7 不同量化向量表征方法在增量索引的时空开销

                    这进一步证明了该方法不仅适用于全量索引的初始构建, 也能够有效支持大规模向量数据的在线扩展与动态
                 更新需求. 我们在原型系统中也展示了这一量化表征方法对更大规模数据的处理能力.
                  5   总 结

                    本研究旨在解决大规模向量数据的索引和查询挑战时所面临的维度和表征问题                             [32] . 本文提出一种量化表征
                 方法——权重残差向量量化 (WRVQ), 平衡了检索的判别能力与存储与计算的时空效率. 相较于近期的量化向量
                 表征方式, 该方法具有更高的存储效率和检索性能.
                    我们的研究结果表明, 在向量数据库的数据存取和知识查询等场景中, 量化表征是一种有效的解决方案, 有望
                 为处理大规模嵌入表示的向量数据提供更可靠的表示、存储和检索途径.
                    我们通过构建面向生成式场景的多模态音视频实时交互等应用场景                         [39,40] , 进一步验证了量化表征的性能优
                 势, 并在大规模文本和多模态等数据集上验证了量化表征的有效性.

                 References
                  [1]   Xiao ST, Liu Z, Han WH, Zhang JJ, Shao YX, Lian DF, Li CZ, Sun H, Deng D, Zhang LJ, Zhang Q, Xie X. Progressively optimized bi-
                     granular document representation for scalable embedding based retrieval. In: Proc. of the 2022 ACM Web Conf. Lyon: ACM, 2022.
                     286–296. [doi: 10.1145/3485447.3511957]
                  [2]   Zhang QX, Xu ST, Chen Q, Sui GX, Xie JD, Cai ZZ, Chen YQ, He YX, Yang YQ, Yang F, Yang M, Zhou LD. VBase: Unifying online
                     vector similarity search and relational queries via relaxed monotonicity. In: Proc. of the 17th USENIX Symp. on Operating Systems
                     Design and Implementation. Boston: USENIX Association, 2023. 377–395.
                  [3]   Ge TZ, He KM, Ke QF, Sun J. Optimized product quantization. IEEE Trans. on Pattern Analysis and Machine Intelligence, 2014, 36(4):
                     744–755. [doi: 10.1109/TPAMI.2013.240]
                  [4]   Mikolov T, Chen K, Corrado G, Dean J. Efficient estimation of word representations in vector space. In: Proc. of the 1st Int’l Conf. on
                     Learning Representations. 2013.
                  [5]   Pennington J, Socher R, Manning C. GloVe: Global vectors for word representation. In: Proc. of the 2014 Conf. on Empirical Methods in
                     Natural Language Processing (EMNLP). Doha: ACL, 2014. 1532–1543. [doi: 10.3115/v1/D14-1162]
                  [6]   Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proc. of the
                     31st Int’l Conf. on Neural Information Processing Systems. Long Beach: Curran Associates Inc., 2017. 6000–6010.
                  [7]   Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional Transformers for language understanding. In: Proc.
                     of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
                     Minneapolis: ACL, 2019. 4171–4186. [doi: 10.18653/v1/N19-1423]
                  [8]   Conneau A, Kiela D, Schwenk H, Barrault L, Bordes A. Supervised learning of universal sentence representations from natural language
   150   151   152   153   154   155   156   157   158   159   160