This dataset is the full Llaza pretraining-data mixture for zip2zip language-model pretraining.
It combines general web text, code, math, and multilingual web text with byte-based top-level mixture ratios.
Domain
Source
Target byte ratio
General
HuggingFaceFW/fineweb-edu, sample-100BT
50%