This is an in-process eval (internally consistent base-vs-tuned), not a public leaderboard harness — the absolute % is not directly comparable to other boards. The delta is the claim. A small model + a quick QLoRA buys format-alignment and a few points, not a new tier; the value is the rigor (measure honestly, align train to eval).
By WITCHEER · rig: github.com/notwitcheer/llm-bench-rig