Next.js Selfbench is a 47-task software-engineering evaluation packaged for the Harbor evaluation framework. Each task asks an agent to implement a change in a frozen revision of vercel/next.js, then evaluates the resulting patch with task-specific tests.
This repository contains the raw evaluation only. It does not include model outputs, scores, costs, or benchmark result artifacts.
Download the raw task package with the Hugging Face CLI:
hf… See the full description on the dataset page:
https://huggingface.co/datasets/dari-ai/nextjs-selfbench.