Addition, subtraction, multiplication, and exact division
One, two, and three operators
Standard operator precedence
Parenthesized two- and three-operator expressions
Zero operands
Boundary and high-carry examples
One- to three-digit operands in the ArithMark-focused release mix
Field
Value
Unique generated examples
644,655
Unique corpus tokens
10,000,014
Full-run packed tokens
1,000,013,824
Full-run supervised answer tokens
246,037,802
Other datasets used
0
The corpus was streamed repeatedly with a 50% one-operator, 30% two-operator, and 20% three-operator training mix.
Training used answer-only supervision. Context and EOS labels were masked, so the objective directly optimized the continuation tokens scored by ArithMark.
Contamination Guard
The training distribution and surface form were deliberately aligned to ArithMark 2.0. To prevent exact benchmark leakage, all 2,500 normalized ArithMark contexts were loaded before generation and hard-excluded.
The generator dropped 135,536 candidate collisions, deduplicated normalized expressions on disk, and asserted zero exact context leakage before writing the final manifest.
Training Setup
Field
Value
Sequence length
256
Micro batch
32
Gradient accumulation
4
Tokens per optimizer step
32,768
Full-run optimizer steps
30,518
Released checkpoint step
29,000
Optimizer
AdamW
Betas
0.9, 0.95
Peak learning rate
1e-3
Warmup steps
100
LR schedule
Linear warmup with cosine decay
Minimum LR ratio
0.3
Weight decay
0.1
Gradient clipping
1.0
Seed
42
Evaluation
Self-reported evaluation on all 2,500 examples from AxiomicLabs/ArithMark-2.0.
The published score uses raw summed log-likelihood over each candidate continuation, matching the benchmark implementation. The release checkpoint was evaluated in float32.
Benchmark Summary
Model
Parameters
ArithMark 2.0
MathBananaMind-1.1
2.90M
90.28%
Score by Operator Count
Operator count
Correct
Total
Accuracy
1
1,139
1,250
91.12%
2
698
750
93.07%
3
420
500
84.00%
Single-Operator Scores
Operator
Topic
Correct
Total
Accuracy
+
Addition
523
538
97.21%
-
Subtraction
411
438
93.84%
*
Multiplication
105
144
72.92%
/
Division
100
130
76.92%
Accuracy on Every Example Containing Each Operator
These groups overlap because a mixed expression can contain multiple operators.
Operator
Correct
Total
Accuracy
+
1,393
1,514
92.01%
-
1,100
1,208
91.06%
*
745
902
82.59%
/
438
487
89.94%
Score by Topic
Topic
Accuracy
Addition
97.21%
Subtraction
93.84%
Multiplication
72.92%
Division
76.92%
Mixed two operators
95.19%
Parentheses with two operators
90.70%
Mixed three operators
88.84%
Parentheses with three operators
79.46%
Repository Files
File
Description
config.json
Transformers configuration for mathbananamind
model.safetensors
Step-29,000 model weights
tokenizer.json
Custom 8k digit-aware tokenizer
tokenizer_config.json
Tokenizer metadata
generation_config.json
Deterministic generation defaults
configuration_mathbananamind.py
Custom Transformers configuration class
modeling_mathbananamind.py
Custom Transformers causal LM implementation
evaluation_results.json
Machine-readable benchmark summary
Usage
This model uses custom architecture code, so load it with trust_remote_code=True.
The model was trained for direct continuations such as expression = answer. It was not trained to follow chat instructions or produce chain-of-thought explanations.
Intended Use
MathBananaMind-1.1 is intended for small-model arithmetic research, continuation-likelihood evaluation, digit-tokenizer experiments, and lightweight local inference.