Raw lm-evaluation-harness result files backing
the rendered comparison table
and the leaderboard Space.
Every number in either of those is read directly from one of the files in
this repo -- nothing here is hand-retyped.
Generated by benchmarks/run_ultimate_comparison.py on 2026-08-17 03:57 UTC. All models below were run through the same lm_eval harness, same task… See the full description on the dataset page:
https://huggingface.co/datasets/DeependraVerma/legal-slm-benchmark-comparison.