Serialized corpus pools for the CTC long-context suite's data generators
(allenai/OLMo-core, branch prasann/ctc,
pip package ctc/). Each file is the output of the one build step that needs heavy machinery — a
GPU cross-encoder, a pyserini/Lucene index, an LLM mining run, or a multi-gigabyte download —
captured once, so that anyone can build train and eval data at any context scale (2k to 10M+
tokens per example) on a bare pip install: no GPU, no Java, no API key.… See the full description on the dataset page:
https://huggingface.co/datasets/PrasannSinghal/ctc-seed-pools.