A benchmark suite of 50 open-source tasks for evaluating AI coding agents on real-world software engineering challenges. This dataset contains 25 integration tasks and 25 observability tasks.
For the evaluation harness (setup, running evaluations, scoring), see:
https://github.com/Mercor-Intelligence/apex-swe
For the technical report, see the arxiv paper:
https://arxiv.org/pdf/2601.08806
For an accessible overview of APEX-SWE, see the blog:… See the full description on the dataset page:
https://huggingface.co/datasets/mercor/APEX-SWE.