The training mixture for the FIRST model-eval-model (self-arm) organism. Pull
mixture.jsonl and train — this is the artifact the 2026-08-07 fine-tune should use.
field
value
experiment
20/80-by-examples SFT mixture: 2,000 model-eval-model self-evaluation docs (m1/m2) + 8,000 spec-filtered Table-2 instruction rows, for LoRA SFT of Qwen3.6-27B (the model-eval-model twin of the self-reflection arm)