Held-out evaluation set used for the hierarchical kill-gate verdict:
Format ≥95% → Safety ≥90% Item-9 sensitivity → Utility ≥10pp Likert
accuracy vs base Gemma 4 E2B. Any single failure → drop fine-tune; ship
base; document honestly.
198 main evaluation rows — stratified random teacher carve-out
(persona × Likert × scale strata mirror the Phase D dad-review sample).
IN-DISTRIBUTION CAVEAT: drawn… See the full description on the dataset page:
https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-evaluation.