This dataset consolidates four synthetic, English, plain-completion tool-use
corpora into one training source. It contains 325,387 documents and
356,011,942 tokens, counted with the pinned
HuggingFaceTB/SmolLM2-135M tokenizer at revision
93efa2f097d58c2a74874c7e644dbc9b0cee75a2.
The data covers broad tool-pattern structure alongside narrower on-device agent
workflows such as calculator, calendar, email, reminders, notes, timers… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-tool-use-mix-356m.