Views
No views yet
| Benchmark | Base | RL (GRPO) | ContextRL (Ours) |
|---|---|---|---|
| SWE-Bench Verified | 5.00 | 6.20 | 7.00 |
| SWE-Bench Lite | 2.70 | 2.70 | 4.00 |
| LiveCodeBench v6 | 44.6 | 46.3 | 47.4 |
| LongBench v2 (Overall) | 31.6 | 31.8 | 33.2 |
| LongBench v2 (Long) | 27.8 | 26.9 | 29.6 |
| NIAH | 98.8 | 98.5 | 99.0 |
transformers.
Training and evaluation code, data construction pipelines, and detailed configurations are
available in the repository:
👉 https://github.com/xupy2003/ContextAwareRL
Please refer to the repo's README for environment setup, inference scripts, and
reproduction instructions.