The 7 problem definitions for KernelBench-Hard, a benchmark for autonomous LLM coding agents writing GPU kernels on a single Blackwell GPU (RTX PRO 6000, sm_120, CUDA 13.2).
Companion datasets:
Infatoshi/kernelbench-hard-runs — 84 agent transcripts, winning solutions, leaderboard, reward-hack annotations
Live site:
https://kernelbench.com/hard
Methodology blog:
https://kernelbench.com/blog/hard
Source repo:… See the full description on the dataset page:
https://huggingface.co/datasets/Infatoshi/kernelbench-hard-problems.