PDL-SWE-Bench is an agentic software-engineering benchmark maintained by
Poindexter Labs — the SWE sibling of
PDL-Bench.
Each task drops an agent into an original, internally-authored code
repository with an engineering issue written as prose, a passing public test
suite, and a fixed token budget. The agent's submitted patch is graded
against a held-out acceptance suite it never saw during the episode.
Like PDL-Bench, this is an open benchmark (HLE-style): the… See the full description on the dataset page:
https://huggingface.co/datasets/Poindexter-Labs/PDL-SWE-Bench.