An ultra-high quality, large-scale dataset specifically designed and mathematically curated for Supervised Fine-Tuning (SFT) of 3-9B parameter class Large Language Models (LLMs). The dataset places a heavy focus on endowing models with "agentic" capabilities, including advanced tool use, deep multi-turn reasoning, intelligent function calling, and structured code execution.
Total Samples: 129,008Splits: Train (119k), Validation… See the full description on the dataset page:
https://huggingface.co/datasets/ankushthakurr09/whiteswan_agentic_A1.