Views
No views yet
Well-engineered 7–8B models achieve 100% output consistency at T=0.0, while a 120B model reaches only 12.5% consistency, regardless of configuration.
| Tier | Models | Consistency @ T=0.0 | Status | Recommended use |
|---|---|---|---|---|
| 1 | Granite-3-8B, Qwen2.5-7B | 100% | ✅ Production-ready | All regulated tasks |
| 2 | Llama-3.3-70B, Mistral-Medium-2505 | 56–100% | ⚠️ Task-specific | SQL / structured only |
| 3 | GPT-OSS-120B | 12.5% | ❌ Non-compliant | Not for compliance |