Views
No views yet
ModernBERT-Small masked language model (RoPE, GeGLU, alternating local/global
attention, 384 hidden size, 16 layers, 6 attention heads) trained from scratch on the
BabyLM 2026 Strict-Small 10M-word corpus, as part of an ablation study on
parameter-efficient token embedding layers for developmentally-plausible pretraining under the
BabyLM Challenge's strict-small compute and data budget.vocab_size x hidden_size, which for this
model's 30,522-token vocabulary and 384 hidden size would be an 11.7M-parameter dense lookup
table. This checkpoint's base representation is the same compositional byte-n-gram embedding
used across this ablation sweep: each vocabulary token is decomposed into its raw UTF-8 byte
sequence, every byte n-gram (n = 1 to 4) is extracted and hashed into one of 8,192 shared
buckets, and the token's representation is the mean of its buckets' embeddings (a
128-dimensional EmbeddingBag table, ~1.0M parameters) rather than a dedicated per-token row.vocab_size x 128 embedding added to the composed vector, scaled per
token by frequency / (frequency + tau) (tau = 100) so common tokens lean more on their own
residual and rare tokens rely mostly on the composition. Critically, this checkpoint is the
shuffled-frequency control: the frequency values used to compute each token's gating weight
are randomly permuted across the vocabulary (fixed seed 42) before being applied, so a token's
residual is gated by an unrelated token's frequency rather than its own. This isolates whether
any benefit from the frequency-gated residual comes from the correct frequency signal, or
merely from having some per-token residual-plus-gating mechanism -- a null-effect control
against the true frequency-gated variant in this sweep.1from transformers import AutoModelForMaskedLM, AutoTokenizer
2
3model = AutoModelForMaskedLM.from_pretrained(
4 "remg1997/modernbert-small-phase6-composition-frequency-shuffled-babylm2026",
5 trust_remote_code=True,
6)
7tokenizer = AutoTokenizer.from_pretrained(
8 "remg1997/modernbert-small-phase6-composition-frequency-shuffled-babylm2026"
9)chck_{N}M branch of this repository corresponds to a BabyLM-Challenge compliance
checkpoint (one per N million words of training data seen); main points at the final,
fully-trained checkpoint.babylm-eval harness
(BLiMP, EWoK, entity tracking, COMPS, Global PIQA, reading-time correlation, GLUE/SuperGLUE
fine-tuning, and Age-of-Acquisition word-surprisal correlation) under the strict-small track.