This is a complete build of an ORPO dataset for the dataset "OpenO1-SFT" located here:
https://huggingface.co/datasets/O1-OPEN/OpenO1-SFT
Using a SmolLM2-360M-Instruct GGUF at 8-bit quantization, I generated "rejected" rows for each of the prompts, and using the original dataset's answers as the "accepted" column.
I chose to use such a small model for 3 reasons:
I am broke but love researching ML and AI.
The time it… See the full description on the dataset page:
https://huggingface.co/datasets/Qurtana/openo1-sft-orpo.