Views
No views yet

| Architecture | Mixture-of-Experts (MoE) |
| Total Parameters | 48B |
| Activated Parameters | 3B |
| Number of Layers (Dense layer included) | 40 |
| Number of Dense Layers | 1 |
| Attention Hidden Dimension | 2048 |
| MoE Hidden Dimension (per Expert) | 768 |
| Number of Attention Heads | 32 |
| Number of Experts | 256 |
| Selected Experts per Token | 8 |
| Number of Shared Experts | 1 |
| Vocabulary Size | 129K |
| Context Length | 128K |
| Attention Mechanism | MLA |
| Activation Function | SwiGLU |
| Benchmark | JoyAI-LLM Flash-base | Qwen3-30B-A3B-base |
|---|---|---|
| MMLU | 84.70 | 82.12 |
| MMLU-Pro | 73.14 | 61.76 |
| CMMLU | 83.09 | 83.60 |
| HumanEval | 85.37 | 87.80 |
| LiveCodeBench | 39.91 | 37.34 |
| GSM8K | 88.78 | 90.37 |
| MATH | 78.16 | 59.60 |
| MATH 500 | 77.00 | 58.00 |