The exact 3.86 GB text mixture used to train
ajaxdavis/alpha-er, a 100M-parameter model trained
end to end on a from-scratch, CUDA-free GPU stack.
This is published as the precise artifact that was trained on, byte for byte, so the run is
reproducible rather than approximately describable.
A single <|end_of_text|>-delimited plain-text file, interleaved from three sources in
proportional blocks so that every source is… See the full description on the dataset page:
https://huggingface.co/datasets/ajaxdavis/alpha-er-corpus.