16 effective layers, trained at a 1e18 FLOP budget at this architecture's
compute-optimal width. One of a six-seed set (seeds 42-47) in which only the
training data order varies; model initialisation is fixed across seeds.
field
value
architecture
moe
d_model
832
d_ff
2240
width_ratio
6.5
base_d_model
128
base_d_ff
384
data seed
42
init seed
42
Loads with trust_remote_code=True. Trained with context length 1024; set
max_length=1024 when evaluating.