M.I.R.O.N. (Multi-aspect Inference Robustness on Objective Next-tokens)
M.I.R.O.N. is a specialized benchmark designed to evaluate the impact of tokenization and architectural constraints on the generation quality of small, Base language models (SLMs).
Unlike global benchmarks (MMLU, GSM8K), MIRON focuses on the atomic capabilities of a model: morphological generalization, noise robustness, and factual integrity within a simple next-token prediction task.
🎯 Main Goal… See the full description on the dataset page: https://huggingface.co/datasets/apsua/MIRON_Benchmark.