This repository contains a subset of the gfissore/arxiv-abstracts-2021 dataset, specifically the first 10,000 samples. It includes embeddings generated from three different models.
id: Unique identifier for each entry.
content: The abstract text from the arXiv paper.
categories: The categories associated with the paper.
embedding: The embedding representation of the… See the full description on the dataset page:
https://huggingface.co/datasets/sondalex/arxiv-abstracts-2021-embeddings-10000.