Views
No views yet
| Benchmark (GENOMA harness) | v1.1L (14B dense) | v1.2L (30B-A3B) |
|---|---|---|
| Coding — v4_hard pass@1, hidden-test sandbox (41 tasks) | 57.7% (N=3) | 73.2% (N=2) |
| Plan-grade orchestration (30 tasks, calibrated LLM-judge) | 0.832 | 0.848 |
| Execution-grade orchestration (failure-handling outcome judge) | 0.437 | 0.482 |
| MMLU-Pro (n=300) | 41.0% | 50.3% |
| Active parameters per token | 14B | 3.3B (~4× cheaper inference) |