This is a multi-model Serbian instruction dataset. It was generated natively and built to SFT the from-scratch Serblo-125M. Every answer was generated directly in Serbian by models a native speaker vetted. Translationese was treated as a defect class, with dedicated filters.
serblo-sft-full.jsonl: 75,452 pairs. The complete release artifact; nothing generated was thrown away that passed cleaning.… See the full description on the dataset page:
https://huggingface.co/datasets/sterlixlol/serblo-sft.