Spider DPO 1040 is a compact Text-to-SQL training dataset for supervised fine-tuning and Direct Preference Optimization. It contains 1,040 preference pairs derived from frontier-model disagreements on Spider V1, plus 7,000 supervised Spider train examples formatted for LLaMA-Factory.
The dataset was created for the companion LoRA adapter jk200201/qwen2.5-coder-7b-sql-dpo.
The DPO preference pairs in this repository were… See the full description on the dataset page:
https://huggingface.co/datasets/jk200201/spider-dpo-1040.