Views
No views yet
Qwen/Qwen3-1.7B with LoRA.| Method | Greedy accuracy | Sampled pass@1 | Sampled pass@4 |
|---|---|---|---|
| Continuous GRPO | 26.5% | 31.0% | 35.5% |
| Fixed staged GRPO | 34.5% | 34.5% | 39.5% |
| LLM controller | 36.5% | 37.5% | 40.5% |
runs/ directory contains metrics, evaluation samples, configuration
history, controller decisions, logs, plots, and all saved LoRA checkpoints.