Views
No views yet
deberta-base-75k-sam,
with extended vision initialization: in addition to 21,134 tokens
seeded from real grounding data (Flickr30k Entities, RefCOCO/g/+, THINGS;
SAM ViT-B features), 737 further tokens covering 1,151 concrete
zero-support words were seeded via a synthetic pipeline (LLM-written
scene descriptions → SDXL-Turbo images ×3 → OWLv2 open-vocabulary
detection → SAM features pooled in detected boxes, through the same
extraction code). Total seeded: 21,871/75,000. The synthetic pipeline
affects embedding initialization only — no synthetic images or
descriptions enter the training text.bb24.train) from our BabyLM
2024 submission (Edman et al. 2024, "Are BabyLMs Second Language
Learners?") — a mixture
of LLM-synthesized paraphrase/contrastive data (SynCSE-partial; Zhang et
al. 2021) and portions of the official BabyLM corpus (Simple Wikipedia,
Gutenberg, Switchboard). Within the strict-small 10M-word budget.chck_1M … chck_100M, step1000 …
step25740; main = final. Code, benchmark, analyses:
https://github.com/bylinina/augustinian_babylm