Reproducible evaluation benchmark for Triton GPU-kernel code generation by LLMs — measures buildability, numerical correctness against a deterministic test suite, and end-to-end speedup vs. a GPU-matched golden reference.
Paper: arXiv:2605.04956 · hf.co/papers/2605.04956
Evaluation harness:
https://github.com/BonnieW05/KernelBenchX
tasks
176
Benchmark task specs + PyTorch reference + deterministic test harness… See the full description on the dataset page:
https://huggingface.co/datasets/BonnieWang/KernelBenchX.