A ready-to-search snapshot of the machine-learning literature on arXiv: 377,093 papers,
28,723,432 full-text chunks, every one embedded with
google/embeddinggemma-300m, plus the
BM25 index, the citation graph, and the raw HTML it was all built from.
This is not a load_dataset() corpus of loose text files. It is the working state of a
retrieval system — a SQLite database and flat memory-mappable vector files — so… See the full description on the dataset page:
https://huggingface.co/datasets/JamesBedichek/lara-corpus-ML-08-17-26.