Views
No views yet

| Metric | Old model | New model (+ synth) |
|---|---|---|
| Overall accuracy | 72.1% | 86.2% |
| Valid rate | 96.6% | 96.8% |
| Accuracy on valid | 74.6% | 89.1% |
| Frequency bucket | Old model | New model |
|---|---|---|
| Very frequent | 80.2% | 86.4% |
| Frequent | 74.4% | 89.4% |
| Medium | 67.0% | 82.9% |
| Rare | 65.4% | 65.4% |
Note on comparison: both models are evaluated on the same synthetic test split. The old model was trained on the original (non-synthetic) dataset; the new model was trained on the synthetic train split. Gains on frequent and medium codes are likely due to the increased training data for those codes, while the lack of improvement on rare codes is expected since not a lot of synthetic data was generated for those.
| Parameter | Value |
|---|---|
| Base model | Qwen3-0.6B |
| Method | GRPO |
| Infrastructure | Jean Zay (IDRIS) |
| Temperature | 1.2 |