The arXiv-model2vec dataset contains embeddings for all arXiv paper abstracts and their corresponding titles, generated using the Model2Vec proton-8M model. This dataset is released to support research projects that require high-quality semantic representations of scientific papers.
Source: The dataset includes embeddings derived from arXiv paper titles and abstracts.
Embedding Model: The embeddings are generated using the Model2Vec… See the full description on the dataset page:
https://huggingface.co/datasets/sleeping-ai/arXiv-abstract-model2vec.