A high-quality subset of R2E-Gym optimized for multi-agent RL training (MAGRPO).
This dataset is filtered for optimal gradient signal in Level 3 (test execution) rewards:
2-20 failing tests per instance (good gradient signal)
Collaboration-suitable (AI-verified two-agent task decomposition)
Test failures in prompt (explicit error context for the model)
Test Count
L3 Reward per Fix
Problem