A Brief Survey of Vector Databases

Краткий обзор векторных баз данных
Hongbin Huang, Xingrui Xie, Han Liu, Wenzhe Hou
2023-12-15

ChromaMilvusPineconehigh-dimensional datasimilarity metricssimilarity search algorithmsvector databases
The explosive growth of massive high-dimensional data requires capabilities for data processing, storing, and analyzing. This brings significant challenges to traditional databases due to the poor ability to handle high-dimensional data and its original design for stand-alone machines. Fortunately, vector databases have provided a practical solution for the management and analysis of high-dimensional data. Especially, they retrieve results related to the query efficiently after encoding various forms of data (e.g., text, image, and video) into vectors. The purpose of this paper is to offer insight into vector databases by presenting a brief survey. Firstly, the workflow of vector databases including indexing and querying, is detailed along with a specific case. Subsequently, we elaborate on the related methods applied in vector databases, which are the core techniques to enhance search efficiency and reduce computational overhead, particularly similarity search algorithms and similarity metrics. Further, we introduce widely used vector database products (e.g., Pinecone, Chroma, and Milvus) and compare them from multiple factors that should be taken into consideration. We also discuss potential avenues for future research in this domain. To conclude, this survey provides a comprehensive understanding of vector databases for retrieval from vast high-dimensional datasets.
1
Core techniques for vector databases include similarity search algorithms and similarity metrics that improve search efficiency and reduce computational overhead.
2
The paper details the workflow of vector databases, including indexing and querying, and provides a specific case illustrating that workflow.
3
The paper identifies potential avenues for future research in vector databases to better support retrieval from vast high-dimensional datasets.
4
The survey compares widely used vector database products (e.g., Pinecone, Chroma, Milvus) across multiple practical factors for selection and use.
5
Vector databases address limitations of traditional databases for massive high-dimensional data by enabling efficient retrieval after encoding diverse data (text, image, video) into vectors.

Vector databases

Their workflow, core methods and techniques for efficient retrieval from high-dimensional data (indexing, querying, similarity search algorithms and metrics), comparison of products, and avenues for future research

Publication Details
Publication Date
2023-12-15
Journal
Publisher
ISSN
Cited by
29
Access Type
Author Information
Authors
Hongbin Huang
Xingrui Xie
Han Liu
Wenzhe Hou
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%