In the VIBE benchmark we consider datasets whose queries are both in-distribution and out-of-distribution.

The following table reports information about the size and dimensionality of each dataset, along with links that allow to download them.

Here below we report the first two PCA components of data and queries for each dataset. Selecting a dataset in the table above allows to update the visualization.

Along with the PCA we display the distribution of Mahalanobis distances between the data points and the data and the query points and the data.

To characterize the difficulty of queries, and hence of the workloads associated with each dataset, we consider the local relative contrast dimension at \(k\).1 Higher values indicate more difficult queries.

The plot below reports the distribution of local relative contrast dimensions at \(k=100\) across the benchmark datasets,2 with datasets arranged in decreasing order of difficulty, top to bottom.

Footnotes

  1. Martin Aumüller and Matteo Ceccarello, “The role of local dimensionality measures in benchmarking nearest neighbor search”, Information Systems 101 (2021), 101807.↩︎

  2. Datasets with inner product similarity are omitted from the plot, as the inner product is not a metric, and the local relative contrast dimension is not well-defined for non-metric distances.↩︎