nano/tokenized_nano.bin
~1GB
预 tokenize 好的 0.5B tokens(uint16, seed=42 从 3.4B 随机采样 ~15%),直接训
from huggingface_hub import snapshot_download
path =… See the full description on the dataset page:
https://huggingface.co/datasets/lilingsunny/llm101-olmo3-zh-demo-data.