A
best-of-N specialist variant of
hemlang/Hemlock-Apothecary-7B:
the epoch-3 checkpoint of an execution-reward GRPO run
(
grimoire /
hemlock-rl)
on
hemlang/hemlock-codex3-SFT.
Mid-training RL dispersed the policy: lower first-try reliability, but the
highest
any-shot solve rate of any Hemlock model measured to date. Use the base Apothecary for
single-shot generation; use this model when sampling several candidates and verifying with
the interpreter — e.g.
hembot --retry — where solve-rate
within k samples is what matters. Notable per-level: systems/concurrency L4 pass@1 4/7 → 5/7,
debugging L6 pass@5 3/5 → 4/5.
GRPO with LoRA r=16, β=0.1 KL, G=8, dynamic sampling, 267 deduplicated verified prompts,
constant lr, checkpoint at epoch 3 of 4 (the epoch-4 endpoint re-concentrates and loses the
pass@5 advantage). Q8_0 GGUF included.