Audited one-pass release SFT corpus for LiLM1. It contains 500M non-padding
training sequence tokens at a maximum length of 4096. The mixture is broad
assistant SFT with a 30% tool/JSON specialization and 5% exact-distribution
pretraining replay.
The tool/JSON category deliberately contains three interfaces: response-masked
ChatML tool episodes, original plain-completion tool workflows used during
pretraining, and valid standalone JSON used… See the full description on the dataset page:
https://huggingface.co/datasets/glouriousgautam/lilm1-release-sft-v1-500m.