A Direct Preference Optimization (DPO) dataset of (prompt, chosen, rejected) triples used in Phase 2 of the COMPASS project to align a Japanese VLM's LLM backbone toward correct mathematical reasoning. The chosen responses are chain-of-thought traces distilled from a Qwen3-30B teacher in the structured
// XML format. The rejected responses are synthetically generated by corrupting the chosen responses under three strategies, mixed… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-reasoning-dpo.