Task suite for RSIBench: A Counterfactual Test of Recursive Self-Improvement
in Coding Agents — 300 self-contained Python patch-and-test tasks
(200 train / 50 validation / 50 test).
Training tasks drive agent self-evolution: the agent runs them, reads
failures, and proposes updates to its own harness source.
Validation tasks gate acceptance of self-proposed harness updates.
Test tasks are held out and used only for final RSI-lift scoring.… See the full description on the dataset page:
https://huggingface.co/datasets/AgPerry/rsi-bench.