[Paper] [Code]
A pre-built Lucene BM25 index over karpathy/fineweb-edu-100b-shuffle—the exact pre-training corpus used by the nanochat family of language models. Built with Anserini.
This index is part of the NanoKnow project: github.com/castorini/NanoKnow