10,000 text samples from HuggingFaceFW/fineweb (train split), streamed and
saved as-is (no shuffling needed beyond the source stream's own order).
Companion single-source dataset to
vuhaian/24_collected
(23,000 examples across 23 sources) — used as the single-source comparison
baseline in continue-pretraining experiments.
fineweb_10k.jsonl -- raw text, one JSON object per line:
{"text": ..., "source": "fineweb", "repo_id":… See the full description on the dataset page:
https://huggingface.co/datasets/vuhaian/fineweb_10k.