| Model Baseline | Direct Code Leakage (↓) | Pedagogical Utility (1-5) (↑) | Conceptual Accuracy % (↑) |
|---|---|---|---|
| Gemini 3.5 Flash (Google Frontier) | 0.0% | 4.79 / 5.0 | 98.7% |
| GPT-5.4-mini (Proprietary) | 0.0% | 4.67 / 5.0 | 98.7% |
| Socratic Muse-30B (SFT+DPO) | 0.0% | 4.75 / 5.0 | 90.0% |
| Socratic Llama-8B (SFT+DPO) | 0.0% | 3.54 / 5.0 | 76.0% |
| Base Llama-3.1-8B-Instruct | 1.3% | 2.55 / 5.0 | 20.0% |
| Qwen2.5-Coder-7B-Instruct | 6.0% | 2.37 / 5.0 | 20.0% |