This dataset contains the embeddings for the TREC-RAG Corpus 2024 embedded with the Cohere Embed V3 English model.
It contains embeddings for 113,520,750 passages, embeddings for 1677 queries from TREC-Deep Learning 2021-2023, as well as top-1000 hits for all queries using a brute-force (flat) index.
We have a pre-build index that only requires 300 MB available at… See the full description on the dataset page:
https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3.