Views
No views yet
| Model | Blind wins | Grounding | Adjudication | Persona |
|---|---|---|---|---|
| pepper-desk-e2b (this, 2B) | 11/12 | 88.6% | 4.67/5 | 3.75/5 |
| Qwen2.5-7B-Instruct-4bit | 1/12 | 77.5% | 2.67 | 1.42 |
| pepper-7b (persona LoRA) | 0/12 | 50.0% | 2.00 | 1.83 |
bench/ in the
repo. The origin story matters: the first Pepper model failed this
benchmark against its own base (38.1% vs 64.5% grounding) — that failure
became the release gate this model had to clear.DESK NOTES: (a private source-weighing analysis) then ON AIR: (the
broadcast). Consumers show or strip the notes; score only the broadcast.bench/README.md
and the repo's gen_eval_v2 harness. Use max_tokens ≥ 500 — tighter caps
truncate her sign-offs (it cost her one judged bundle).google/gemma-4-e2b-it via mlx-community/gemma-4-e2b-it-4bit