Views
No views yet
deepseek-ai/DeepSeek-R1-0528-Qwen3-8B.
Designed as a tutor that scaffolds for students with diverse learning
disabilities (ID, ASD, ADHD, EBD, SLD-Reading, SLD-Math).1models:
2 - model: OpenLearnLM/special-r1-deepseek-qwen3-8b-think-reward
3 parameters: { density: 0.5, weight: 0.5 }
4 - model: OpenLearnLM/special-r1-deepseek-qwen3-8b-sped-adaptive-think-reward
5 parameters: { density: 0.5, weight: 0.5 }
6merge_method: dare_ties
7base_model: deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
8parameters:
9 normalize: true
10 int8_mask: true
11dtype: bfloat16
12tokenizer_source: unionk=8 student samples for solve-rate
measurement, helpfulness scored by gpt-4o-mini against the 5-criteria
SpEd rubric (scaffolding / language / tone / pacing / disability-appropriate).
Student model: meta-llama/llama-3.1-8b-instruct.| Metric | DARE-TIES v2 | Parent #3 (think-reward) | Parent #7 (sped-adaptive) |
|---|---|---|---|
| Pre-dialog solve rate | 0.157 | 0.150¹ | 0.200¹ |
| Post-dialog solve rate | 0.472 | 0.512¹ | 0.400¹ |
| Δ Solve Rate | +0.315 | +0.362¹ | +0.200¹ |
| Helpfulness mean | 0.837 | 0.800¹ | 0.820¹ |
| scaffolding pass | 0.836 | 0.900¹ | 0.800¹ |
| language pass | 0.784 | 0.700¹ | 0.800¹ |
| tone pass | 0.992 | 1.000¹ | 0.900¹ |
| pacing pass | 0.784 | 0.700¹ | 0.800¹ |
| disability_appropriate pass | 0.788 | 0.700¹ | 0.800¹ |
| Thinking dialog rate | 0.660 | 0.800¹ | 0.800¹ |
| Leak rate (regex) | 0.786 | 1.000¹ | 1.000¹ |
SLERP t=0.5,
TIES density=0.5, DARE-TIES density=0.5, plus weight-ratio
variants dare-ties 0.7/0.3 and dare-ties 0.3/0.7). DARE-TIES v2 at
default 0.5/0.5 was the only candidate that strictly passed the
pre-registered 3-axis decision rule (Δ Solve ≥ max-parent − 0.5pp;
Helpfulness ≥ max-parent; Scaffolding-pass ≥ min-parent).1@misc{openlearnlm_special_r1_dare_v2,
2 title = {special-r1-deepseek-qwen3-8b-merged-dare-v2},
3 author = {OpenLearnLM},
4 year = {2026},
5 howpublished = {\url{https://huggingface.co/OpenLearnLM/special-r1-deepseek-qwen3-8b-merged-dare-v2}},
6 note = {DARE-TIES merge of two GRPO-trained SpEd tutoring models}
7}