Router activation data captured from 20 checkpoints spanning a 100B-token sparse
Mixture-of-Experts pretraining run, with fully documented data provenance for
every input token. Base model: harims95/hobbylm-1b-hf.
For each of 20 training checkpoints, the model was run over an identical frozen
probe set and its router's per-token, per-layer decisions were recorded: which
experts were selected, the router's raw scores before… See the full description on the dataset page:
https://huggingface.co/datasets/harims95/hobbylm-routing-dynamics.