Feature-addition benchmark for LLMs and coding agents, evaluated against the
Scorch codebase. Each task asks a model
to add a feature (or otherwise extend functionality) to Scorch. Success is
defined as the full pytest suite (original + any new tests the model adds)
passing after the patch is applied inside a Docker container.
This is the dataset artifact for the TensorBench paper (NeurIPS 2026
Evaluations & Datasets track, double-blind submission).
At a… See the full description on the dataset page: https://huggingface.co/datasets/tensorbench/tensorbench-1.0.