This checkpoint is the validity-reward GRPO baseline.
Each completion is rewarded for producing a hypothesis that is consistent with
the visible evidence; the reward has no explicit set-diversity term.
Exact training and evaluation configurations, per-file model hashes, and the
pinned dataset revision are recorded in release_manifest.json and
provenance/configs/.