FineWeb-style Parquet shards (LLM-training friendly) created by merging multiple Wikipedia language datasets into a single dataset to reduce looping during training.
Format: nanochat-parquet-v1
Layout: shard_*.parquet + metadata.json
Text column: text
Parquet settings: zstd (level 3), row_group_size=1024, use_dictionary=False, write_statistics=False
This mix was built from the already-exported wiki datasets in… See the full description on the dataset page:
https://huggingface.co/datasets/JayJayThrowThrow/europe-wiki-mix.