btprop-rlv3-qwen3-8b
RL v3. Archived negative result.
This is an archived negative result. It is published so the claim it supports can be checked, not because the model is worth deploying -- read the numbers below before using it for anything.
GRPO variant with the prior_wrong denominator clipped in the reward. Trained on wiki-23
100-word-slice evidence, where the perturbation layer has little to work with; superseded by the
wiki-2026 run. Archived for provenance, not for use.
Evaluation protocol
Metrics are per-statement hallucination detection on the BTProp stop-node test split (6 datasets, n=2,225 shared claims), with retrieval, judging and aggregation held fixed so that only the generator differs. See EXPERIMENTS.md in the code repo for the full provenance table.
Code