Views
No views yet
| Benchmark | Metric | zoof-394M (v1.2.2) | SmolLM-360M | SmolLM2-360M | Qwen2.5-0.5B |
|---|---|---|---|---|---|
| Training Tokens | Data Efficiency | 79B | 600B | 4T | 18T |
| PIQA | Physical Commonsense | 69.5 | 71.6 | 71.7 | 69.9 |
| BoolQ | Boolean Reasoning | 59.9 | - | - | - |
| WinoGrade | Pronoun Resolution | 53.8 | 52.8 | 52.5 | 54.1 |
| HellaSwag | Commonsense NLI | 47.0 | 51.8 | 54.5 | 51.2 |
| OBQA | OpenBookQA | 37.2 | 37.2 | 37.4 | 37.4 |
| ARC-E | Science (Easy) | 44.3 | - | - | - |
| ARC-C | Science (Challenge) | 32.3 | - | - | 35.6 |
| SIQA | Social Commonsense | 40.3 | - | - | - |
| MMLU cloze | General Knowledge | 28.5 | 34.4 | 35.8 | 33.7 |
| MMLU | General Knowledge | 29.6 | - | - | - |
| RACE | Reading Comprehension | 38.3 | - | - | - |
Note: Zoof achieves these scores with ~2% of the training compute used for SmolLM2 (79B vs 4T tokens), highlighting the efficiency of the FineWeb-Edu dataset.
RMSNorm (Pre-Norm).F.scaled_dot_product_attention.