A 200 shards subset of karpathy/climbmix-400b-shuffle dataset (Nvidia ClimbMix) of web documents with added pre-computed embeddings and classified topics and formats.
Parquet files keep the nanochat compatible format (row groups, 'text' column), so this can be used as a drop-in replacement of the Karpathy's mix in the nanochat project, where the additional metadata can be used in the code.
Size: 200 parquet shards (~86K rows each, ~16.9M… See the full description on the dataset page:
https://huggingface.co/datasets/ddudek/nanochat-climbmix-annotated.