This dataset contains benchmark evaluation results for a single selected checkpoint from the MyAwesomeModel training run.
Selected checkpoint: step_1000
Eval accuracy from checkpoint config:
The table below lists each benchmark and the score produced by running evaluation/eval.py on the selected checkpoint. Scores are shown with three decimal places.
Math Reasoning: 0.550
Logical Reasoning: 0.819
Common Sense: 0.700… See the full description on the dataset page:
https://huggingface.co/datasets/FuryAssassin/BenchmarkResults-Migration.