Views
No views yet
model=d26_r10), sequence length 2048,
9,184,215,040-token stream; trained on 8×H200 on bulbasaur.41de86425450676dc4d5702fd2955d8fd734331a, config
conf/data/pirate2x2_dose40.yaml (export/conversion code ran at commit
35bf062c0d74448cf0a8a27d655f1d3a3551e24e), arm hash a0e1a0c3a59c, checkpoint step
8,758.ppriors/hf_export/convert.py (bf16 safetensors, custom
trust_remote_code modeling files). Logit/tokenizer/bpb/KV-cache
equivalence against the nanochat checkpoint verified on GPU: logit max abs
diff 0.00e+00; converted val bpb 0.720066 (training-time record
0.720055). Results in verify_results.json, uploaded alongside the
model on HF.trust_remote_code=True. Experiment registry: exp-074
(pretraining-priors project). Sibling instruction-SFT model:
jkminder/pretraining-priors-pirate2x2-d26-dose40-sft.