ArithMark 3.0 is our 3rd gen benchmark for evaluating arithmetic ability in language models. As a changeup from our previous two entries, problems are expressed as continuation style short English word problems rather than bare equations.
The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.
The included… See the full description on the dataset page:
https://huggingface.co/datasets/AxiomicLabs/Arithmark-3.0.