VIBE: Vector Index Benchmark for Embeddings.
- Elias Jääsaari,
- Ville Hyvönen,
- Matteo Ceccarello,
- Teemu Roos,
- University of Helsinki,
- Aalto University,
- ,
- ,
Research Output:
Working paper
Preprint
Open access
Publication Information
Output type
Research Output:
Working paper
Preprint
Original language
EnglishPublication milestones
- Published - 23/05/2025
Publication status
Published - 23/05/2025
Volume
abs/2505.17810Book series
- Book series name: CoRR
ISSN: 0000-0000
Abstract
Approximate nearest neighbor (ANN) search is a performance-critical component
of many machine learning pipelines. Rigorous benchmarking is essential for evaluating the performance of vector indexes for ANN search. However, the datasets of
the existing benchmarks are no longer representative of the current applications of
ANN search. Hence, there is an urgent need for an up-to-date set of benchmarks.
To this end, we introduce Vector Index Benchmark for Embeddings (VIBE), an
open source project for benchmarking ANN algorithms. VIBE contains a pipeline
for creating benchmark datasets using dense embedding models characteristic of
modern applications, such as retrieval-augmented generation (RAG). To replicate
real-world workloads, we also include out-of-distribution (OOD) datasets where
the queries and the corpus are drawn from different distributions. We use VIBE
to conduct a comprehensive evaluation of SOTA vector indexes, benchmarking 21
implementations on 12 in-distribution and 6 out-of-distribution datasets
of many machine learning pipelines. Rigorous benchmarking is essential for evaluating the performance of vector indexes for ANN search. However, the datasets of
the existing benchmarks are no longer representative of the current applications of
ANN search. Hence, there is an urgent need for an up-to-date set of benchmarks.
To this end, we introduce Vector Index Benchmark for Embeddings (VIBE), an
open source project for benchmarking ANN algorithms. VIBE contains a pipeline
for creating benchmark datasets using dense embedding models characteristic of
modern applications, such as retrieval-augmented generation (RAG). To replicate
real-world workloads, we also include out-of-distribution (OOD) datasets where
the queries and the corpus are drawn from different distributions. We use VIBE
to conduct a comprehensive evaluation of SOTA vector indexes, benchmarking 21
implementations on 12 in-distribution and 6 out-of-distribution datasets
Funding Details
This work has been supported by the Research Council of Finland (grant #361902 and the Flagship programme: Finnish Center for Artificial Intelligence FCAI) and the Jane and Aatos Erkko Foundation (BioDesign project, grant #7001702). Martin Aumüller received funding from the Innovation Fund Denmark for the project DIREC (9142-00001B). The authors acknowledge the research environment provided by ELLIS Institute Finland, as well as CSC – IT Center for Science, Finland, for computational resources.
FundersFunding numbers
Research Council of Finland
361902
FCAI
-BioDesign
7001702
DIREC
9142-00001B
Access to documents
License:Other
