DiscoveryBench ctxgraph-8b-dpo Qwen3-8B + DPO (LoRA r16, 147 synth pairs, 3 epochs, beta 0.1) eval 1/2 on 239 REAL tasks; strict 0.0778, answered 156, zero truncation-retry prompts, vista job 933234. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 156/239 answered, mean HMS 0.1193 over answered / 0.0778 strict-239. Part of DPO data generation: 8B strict scores 0.0846/0.0778/0.0832 (above all 30B ctxgraph runs… See the full description on the dataset page:
https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v1.