A large-scale synthetic coding dataset designed for training and fine-tuning 2B parameter language models. Contains 1.5M instruction-response pairs spanning 22 programming languages and 10 task categories, formatted in the ShareGPT/Alpaca hybrid conversation schema.
Schema
ShareGPT/Alpaca hybrid (conversations… See the full description on the dataset page:
https://huggingface.co/datasets/ArcOffical/PiCo-DATA-code.