Distill-SFT on Coder-30B passing patches under the same anchored context used by the student.
KL-regularized GRPO continuation.
Best checkpoint selected by held-out greedy reward.
Measured held-out result from runs/passk_q3_8b_distill_passk/summary.json:
split
greedy pass@1
pass@8
overall
6/35
8/35
single-file
6/18
8/18
multi-file
0/17
0/17
Best checkpoint selection:
field
value
best step
75
held-out greedy
6/35
held-out reward
0.3806
Interpretation: the anchored teacher distill path was non-harmful and improved measured pass@8 by one over the
robust Qwen3-8B KL-GRPO baseline (7/35 -> 8/35). The extra pass@8 coverage is single-file tail coverage, not
multi-file transfer. Treat 8/35 as measured but not seed-checked.
This is a PEFT LoRA adapter, not a standalone full-precision base model.
Model Details
Model Description
Developed by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Model type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More Information Needed]
Downstream Use [optional]
[More Information Needed]
Out-of-Scope Use
[More Information Needed]
Bias, Risks, and Limitations
[More Information Needed]
Recommendations
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.