Reproducible code-LLM evaluation forged on idle TU Delft DAIC RTX Pro 6000 (Blackwell, sm_120) GPUs. vLLM 0.23 (apptainer), bf16, greedy temp=0, pass@1 with bootstrap-over-problems 95% CI. HumanEval+/MBPP+ via EvalPlus (pass_at_1=plus set); CRUXEval-I/-O = input/output prediction, Direct (no-CoT), 2-shot, execution-scored (comparable to the paper's Direct column).
Honest caveat: HumanEval+/MBPP+/CRUXEval are likely in pretraining corpora ->… See the full description on the dataset page:
https://huggingface.co/datasets/D4vidHuang/benchforge-leaderboard.