This is the sentence-transformers/msmarco-corpus dataset,
embedded with Alibaba-NLP/gte-modernbert-base.
For each example, we embed the text directly (no additional instruction prompt).
Embeddings have dimensionality 768.
These embeddings are intended for tasks like large-scale distillation, retrieval, and similarity search.
Because the raw text may exceed the model’s limit, we recommend… See the full description on the dataset page:
https://huggingface.co/datasets/stephantulkens/msmarco-gte-modernbert-pooled.