Pre-built JSONL training splits for the Marco-MoE multilingual expansion experiments.
Contents
splits/ — document-budget tiers (~15 GB)
Files: {lang}_{50k|100k|500k}.jsonl, {lang}_manifest.json… See the full description on the dataset page:
https://huggingface.co/datasets/LiangYan3612/c4.