Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
cayley-10b-k8-3l-mlp_in – AI Model by markhenry | AlphaNeural AI
You can deploy this model and start earning money today!
markhenry
/
cayley-10b-k8-3l-mlp_in
like
0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
cayley-10b-halfk
CayleySAE GPT trained with
half-k
topology: k=8/16/32 per level instead of the standard k=16/32/64.
Architecture
12 layers, 8 heads, d=1024 (~205M params)
CayleySAE at mlp_in: L0 (1024, k=8) → L1 (8192, k=16) → L2 (65536, k=32)
Trained on FineWeb-Edu-10B for 16k iters
Results
Best val loss: 3.1816
Compare: cayley-10b (standard k) val loss 3.173