16 effective layers, trained at a 1e18 FLOP budget at this architecture's
compute-optimal width. One of a six-seed set (seeds 42-47) in which only the
training data order varies; model initialisation is fixed across seeds.
field
value
architecture
base
d_model
896
d_ff
2432
width_ratio
7.0
base_d_model
128
base_d_ff
384
data seed
44
init seed
42
Loads with trust_remote_code=True. Trained with context length 1024; set
max_length=1024 when evaluating.