The benchmark for modern Next.js code generation and completion.
NextBench measures how well a language model can complete real-world Next.js / React / TypeScript code. Every task is an autocomplete prompt — a partial file with the cursor at the end — graded against deterministic checks: must-contain patterns, forbidden patterns, regex matches, and output length.
443 tasks across 16 categories (v0.2)
12-model leaderboard spanning 1.3B–30B parameters; reproducible from… See the full description on the dataset page:
https://huggingface.co/datasets/baablabs/nextbench.