A student scorer from MIRA (Mid-training Rubric Anchoring for Source-Aware Data Selection), fine-tuned to score HTML / front-end QA along a group-specific set of anchor rubric dimensions.
MIRA is a source-aware data selection framework for heterogeneous mid-training corpora. Instead of applying a single global quality rubric, MIRA (1) clusters sources into capability-coherent groups, (2) lets a frontier teacher (Kimi-K2.6) freely propose rubric dimensions and anchors them per group, (3) distills the anchored teacher into a lightweight per-group student scorer, and (4) applies reliability-aware aggregation with per-source retention thresholds.
This repository is one of those student scorers — variant 5 in the QA family, specialized for HTML / front-end QA. Given an in-distribution record, it produces a numerical score and a short rationale for every anchor dimension in this group's rubric.
Model summary
Architecture
Mixture-of-Experts decoder (35B total / ≈3B active params)
Full-parameter SFT on Kimi-K2.6 anchored teacher labels
Domain
HTML / CSS / JS QA — html_linxi (8 sub-domains: game, widget, visual, svg, etc.) and html_sft from the artifact-web/sft corpus. Emphasizes visual quality, interactivity correctness, and browser compatibility.
Anchor rubric
15 group-specific dimensions (group_4_dim_anchors.jsonl in the project repo)
Source count
2 qa sources
Output
Structured (score, rationale) per anchor dimension
Precision
BF16
License
Apache-2.0 (inherits from Qwen3)
Sources covered
This scorer is calibrated for the following mid-training sources in the QA / HTML and front-end group:
Source
Description
html_linxi
HTML QA across 8 sub-domains (game/widget/visual/svg/…) — qa_with_think
html_sft
artifact-web/sft HTML examples
The full source-grouping report (KMeans k=4 / 5 clusters, intra-group cosine similarities) is in the project repo.
Anchor dimensions (15 slots)
The scoring rubric for this group, discovered via Kimi-K2.6 free-form judging and clustered into 15 anchor dimensions (KMeans k=15 over the group's dim-score embeddings). Dimensions below are sorted by cluster size — larger clusters dominate the corpus and carry more signal. Anchor names are read verbatim from this group's group_4_dim_anchors.jsonl; some names recur across slots because semantically related but distinct rubric facets were clustered separately by the teacher.
Slot
Dimension
Cluster size
A1
Educational Value
4,741
A2
Actionability
3,630
A3
Pedagogical Value
3,224
A4
Think-Response Consistency
3,121
A5
Educational Value
2,837
A6
Difficulty Appropriateness
2,634
A7
Problem Clarity
2,469
A8
Think Depth Matches Difficulty
2,347
A9
Technical Precision
2,132
A10
Problem Clarity
1,809
A11
Verifiability
1,595
A12
Efficiency
1,527
A13
Logical Flow
1,014
A14
Training Signal Quality
985
A15
Self-Correction
935
The scorer outputs one [Ai] <dimension>: <score>/10 — <rationale> line per slot, plus overall, training_recommendation, domain_tag, and brief.
MIRA-QA-Group5 lives in Stage 2: it scores the full QA / HTML and front-end corpus so that downstream stages can apply reliability masking and source-aware retention.
Intended use
Primary: Score HTML / front-end QA on this group's anchor dimensions to drive source-aware data selection and filtering.
Secondary: Research on rubric distillation, semantic quality scoring, and reliability diagnostics for heterogeneous training corpora.
Not intended for:
General-purpose chat or instruction following — fine-tuned to emit structured scores, not freeform dialogue.
Single-shot quality judgments without the anchor-dimension prompt template — outputs will be miscalibrated.
Records outside the QA / HTML and front-end group; use the matching sibling scorer instead.
Deployment
The scorer is designed to be served via vLLM behind an OpenAI-compatible endpoint and called in batch from the MIRA scoring pipeline.
The system prompt embeds the top-12 anchor calibration references (canonical examples from clustering) so the student matches the teacher's scoring scale. The full prompt builder, anchor JSONL files, and output parser are in the project repo's scoring/score_qa_anchored.py.
Training details
Teacher
Kimi-K2.6 (free-form rubric discovery in Phase 1; anchored re-scoring in Phase 2)
Training data
Kimi-K2.6 anchored labels on this group's Phase-2 corpus, split into a distillation set + a held-out validation split for reliability diagnostics
Loss
Standard next-token CE over (score, rationale) labels for every anchor dimension
Hyperparameters
Held constant across all MIRA student scorers; full settings in paper Appendix A.4
Validation
Per-dimension teacher–student MAE and Spearman ρ on a held-out split; dimensions failing reliability thresholds are masked post-hoc (Figure 3 in the paper)
Training loss / step curve is preserved in trainer_state.json for full reproducibility.
Headline results (from the paper)
End-to-end downstream evaluation: Qwen2.5-Coder-14B mid-trained on 25B-token MIRA-selected subsets vs. baselines, then SFT, evaluated on 9 code benchmarks across 4 categories.
Method
Code Gen
MultiplE
SQL (EX)
SWE-Multi
Macro Avg
Base + SFT (no mid)
53.91
72.57
64.24
3.67
48.60
Raw Mixture (50B)
53.71
67.42
94.18
40.00
63.83
Random (25B)
52.71
71.44
91.03
35.00
63.23
DataMan (25B)
53.82
71.38
93.84
33.00
63.01
DSIR (25B)
48.74
67.26
95.20
27.00
59.55
PPL (25B)
50.52
57.74
90.66
20.00
54.73
MIRA-Global (25B)
53.12
67.84
94.26
32.00
61.81
MIRA-Group (25B)
54.53
71.85
94.08
36.33
64.20
MIRA-Source (25B)
54.18
72.84
94.38
30.33
62.93
MIRA-Group matches the full 50B-token raw mixture while using only half the tokens, and out-performs all 25B-token selection baselines on the macro average. This scorer is one of the 12 student models used by the MIRA-Group variant.
Sibling models
MIRA releases one student scorer per source-group variant. Use the matching scorer for each record's format:
MIRA addresses source-aware filtering only. Source discovery, mixture-ratio design, curriculum scheduling, deduplication and contamination control remain orthogonal concerns.
This scorer is calibrated against the QA / HTML and front-end group; cross-domain transfer is not advised — use the matching sibling for other source formats.
Some anchor dimensions exhibit high teacher–student MAE and are masked post-hoc during aggregation (see paper §3.4). The model still emits scores for masked dimensions; downstream consumers should re-apply the reliability mask from the project repository.
Calibrated on 2 sources within this group; behavior on out-of-distribution formats is unverified.
Citation
bibtex
1@inproceedings{wang2026mira,
2 title = {MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection},
3 author = {Wang, Haowen and Du, Yaxin and Yang, Jian and Wu, Jiajun and
4 Liu, Shukai and Zhang, Yuxuan and Wang, Pingjie and Chen, Siheng and
5 Zheng, Tuney and Zhou, Ming and Liu, Xianglong},
6 booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
7 year = {2026}
8}