This dataset contains the sparse retrieval index generated by the EviGraph-R indexing pipeline.
It is exported from the finalized shard records after the collection has been written to Qdrant, so the Hub copy matches the indexed corpus that was prepared for retrieval.
One row per indexed chunk.
Original chunk payload metadata used by retrieval and analysis.
Vector columns: sparse_indices / sparse_values.
Source collection:… See the full description on the dataset page:
https://huggingface.co/datasets/lostelf/arxiv_sparse_sample.