Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
The Frontier AI Infrastructure Benchmark for evaluating frontier LLMs and autonomous coding agents on real-world ML systems and LLM infrastructure engineering — also written FAI-Bench / ΦBench.
85 open-source LLM-infrastructure engineering tasks — build a public Docker image, solve the task offline, and score against a grader shipped with the package.