This repository contains a deterministic tokenized representation of a
1,749,978,785,641-token subset of
Zyda-2 for language-model
pretraining.
Files under data/ contain contiguous little-endian uint16 token IDs.
Concatenate shards in numeric order to reproduce the original stream. The
shared 64K BPE tokenizer is stored under tokenizer/tokenizer.json. Shard
manifests provide byte offsets, token counts, and SHA-256 checksums.
The… See the full description on the dataset page:
https://huggingface.co/datasets/zee-drytis/damr-zyda2-64k-tokenized.