
The set of grafted source models, the number of evolution generations, the breeding strategy, dataset composition, and training configuration are proprietary and not disclosed.
| Item | Value |
|---|---|
| Parameters | ~31B (dense) |
| Modality | Text + Image (multimodal) |
| Context length | up to 256K tokens |
| Base family | gemma4-31b (Gemma-compatible architecture) |
| Focus | Administrative & public-sector AI services |
| Method (test-time compute) | Score |
|---|---|
| maj@8 + tie-retry + DELPHI + near-miss maj@32-64 (weighted vote) | 84.34% (167/198) |
| # | Dataset Name | AIHub Link |
|---|---|---|
| 1 | Medical and Legal Professional Books Corpus | 71487 |
| 2 | Financial and Legal Document Machine Reading Comprehension | 71610 |
| 3 | Large-scale Web-based Korean Corpus | 624 |
| 4 | Large-scale Book-based Korean Corpus | 653 |
| 5 | National Records Large-scale AI Learning Corpus | 71788 |
| 6 | Korean Generation-based Common Sense Reasoning Dataset | 459 |
| 7 | Multi-session Dialogue Corpus | pkg1 |
| 8 | Essential Medical Knowledge Data (142GB) | 71875 |
| 9 | Specialized Medical Knowledge Data (206GB) | 71874 |
| 10 | Korean Dialogue Dataset | 272 |
All datasets are publicly available via AIHub (registration required).