This repository contains 46,892,351,992 pretraining tokens derived from
FineMath and
OpenWebMath.
Files under data/ contain contiguous little-endian uint16 token IDs.
Concatenate shards in numeric order to reproduce the original stream. The
shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard
manifests provide byte offsets, token counts, and SHA-256 checksums.
The intended production mixture uses Zyda2 at weight… See the full description on the dataset page:
https://huggingface.co/datasets/zee-drytis/damr-math-64k-tokenized.