This model improves the instruction-following capabilities of Qwen-2.5-3B-Instruct via preference tuning on the
WildChecklists dataset. Full details are provided in Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning.
We report performance on instruction-following and general-chat benchmarks, using GPT-5-mini as the judge. Additional evaluation details and settings are provided in the paper.
Alpacaeval/Arena-Hard: