This dataset stores raw benchmark result artifacts generated by HAKARI-Bench.
Raw results: per-task JSON (.xz) result files measured by HAKARI-bench.
Leaderboard:
https://huggingface.co/spaces/hakari-bench/leaderboard
GitHub repository:
https://github.com/hakari-bench/hakari-bench
Contributing official model results: follow the new model evaluation workflow to evaluate a model and submit results for HAKARI-Bench review:… See the full description on the dataset page:
https://huggingface.co/datasets/hakari-bench/results.