Dataset Card — German Fluency Preference Dataset
Overview
This dataset contains 17,221 German-language preference pairs (chosen/rejected) with chain-of-thought reasoning, assembled and quality-repaired from a multilingual pipeline targeting German translation of the Soofi-10B SFT corpus.
All records carry qwen_confidence: high, meaning only repairs judged high-confidence by the Qwen repair model were accepted.