Per-token multi-vector embeddings for the full MS MARCO v2 passage corpus
(~138.4M passages) plus the dev / dev2 queries, produced with the official
ColBERTv2 checkpoint (colbert-ir/colbertv2.0).
Precision: fp32 (NO residual quantization, NO pooling)
Dim: 128 per token
Corpus token vectors: ~9.41B (avg ~68 tokens/passage)
Total corpus size: ~4.82 TB
These embeddings are uncompressed on purpose (research on… See the full description on the dataset page:
https://huggingface.co/datasets/yaooooo233/msmarco-v2-colbertv2-fp32.