The standing conflict-slice metric of the reuse-vs-recraft programme
(program-wide fix, 2026-08-16). 150 items: for each of three benchmarks,
25 repair-side items the untrained model (Qwen2.5-VL-7B-Instruct,
greedy, 48 new tokens) answers WRONGLY — it follows the misleading text —
and 25 cost-side items it answers CORRECTLY — it resists. A trained
model is scored on both sides:
fixed — repair-side items it… See the full description on the dataset page:
https://huggingface.co/datasets/saketh-chervu/reuse-vs-recraft-slice150.