SpaceV, initially published by Microsoft, is arguably the best dataset for large-scale Vector Search benchmarks.
It's large enough to stress-test indexing engines running across hundreds of CPU or GPU cores, significantly larger than the traditional Big-ANN, which generally operates on just 10 million vectors.
It provides vectors in an 8-bit integral form, empirically optimal for large-scale Information Retrieval and Recommender Systems, capable of leveraging… See the full description on the dataset page:
https://huggingface.co/datasets/unum-cloud/ann-spacev-100m.