Views
No views yet
Qwen/Qwen3-1.7B on 195 rows of
kelsbeans/digestive-coach-dataset
(nested subset of one shuffle; fixed family-disjoint valid set of 81).
Merged fp16 — loads with plain transformers, no peft. Adapter: kelsbeans/qwen3-1.7b-digestive-coach-n195-adapter.c10602623262a8d87b6e468257c37b0647256d5c75863851837a7cc853b7052704e8ef7b6f7f8690| metric | base Qwen3-1.7B | this model |
|---|---|---|
| parse rate | 12.5% | 97.5% |
| schema validity | 2.5% | 97.5% |
| spec adherence | 0.0% | 10.0% |
| field accuracy | 26.7% | 27.4% |
python eval.py --model kelsbeans/qwen3-1.7b-digestive-coach-n195 --base Qwen/Qwen3-1.7B --eval-set eval_scenarios.jsonl --max-new-tokens 1600
(harness in the project repo at the pinned eval-code hash; raw per-example
transcripts ship as JSONL in runs/).