Teacher-feature distillation corpus for the Ogma-v2 10B run. Each row is a
FineWeb-Edu document paired with a 256-dim
Qwen3-Embedding-0.6B
teacher embedding (Matryoshka Representation Learning, truncated to the first
256 dims and stored as float16).
split
rows
notes
train
13,860,000
merged corpus minus the two 70k val slices
val_uniform
70,000
uniform-random sample over the full 14M corpus, seed 20260709