Qwen3-4B IF-RLVR subset4k N=3 outer-CI checkpoints
This public repository contains BF16 Hugging Face exports from one of two
five-epoch Qwen3-4B thinking-model experiments:
- 95% two-sided Student-t outer interval anchors, targeting N=3.
- 99% two-sided Student-t outer interval anchors, targeting N=3.
The pinned best-available data contains all three independent draws for 3,865
of 4,096 prompts, two draws for 63 prompts, and the verified run-1 draw for 168
prompts. No missing draw is imputed, duplicated, or fabricated. N=3 rows use a
Student-t interval with df=2, N=2 rows use df=1, and N=1 rows retain the run-1
point anchor. Every row records its effective sample count and available draws
in the derived cache provenance.
All draws use temperature 1.0, top-p 0.95, top-k 20, and presence penalty 0.0.
The lower reward threshold is the lower endpoint for x-only per-token mean NLL;
the upper threshold is the upper endpoint for x+c per-token mean NLL. Response
IDs/token counts come from draw 1; NLL and PPL fields encode the selected
endpoints.
Training otherwise matches the subset4k A3 strict baseline: Qwen3-4B with
thinking enabled, 4,096 rows, batch 256, 16 steps/epoch, 5 epochs, LR 1e-6,
8 rollouts, 2,048 prompt tokens, 8,192 response tokens, hard-zero lower floor,
and +0.1 reward inside the valid anchor interval. Checkpoints are exported at
steps 16, 32, 48, 64, and 80.
Each checkpoint lives under <experiment>/global_step_<step>/. The
corresponding _manifests/ entry records file sizes, BF16 tensor validation,
and the verified Hub commit.
Pinned additional draws: