For details on how this dataset was used (pre-training pipeline, block packing, GPU benchmarks), see the full write-up: TinyQwen: Understanding the Pipeline for Training an LLM from Scratch
This dataset is the tokenized version of Aquiles-ai/TinyQwen-Data.
Qwen/Qwen3.5-0.8B was used to tokenize it.