hemlang/Hemlock2-Coder-7B improved with
execution-reward GRPO (
grimoire ≥ 2.0 /
hemlock-rl). Completions were executed in the
Hemlock interpreter's sandbox and rewarded for exactly reproducing verified reference stdout.
Strictly improves the base model on
hembench
(zero-shot, n=5, benchmark-overlapping training tasks held out):
Largest gains in syntax (L1 pass@1 2/9 → 4/9, pass@5 7/9 → 8/9) and systems/concurrency
(L4 pass@1 1/7 → 3/7).
Q8_0 GGUF included.