Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
babylm-llama-ar-v1-100Mtokens – AI Model by kkkkkkshy | AlphaNeural AI
You can deploy this model and start earning money today!
kkkkkkshy
/
babylm-llama-ar-v1-100Mtokens
like
0
pytorch
babyllama_kimi
babylm
multilingual
attention-residual
custom_code
en
nl
zh
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
BabyLM 2026 Multilingual: AttentionResidualLlama (v1, 100M tokens)
LLaMA-derived model with Attention Residual mechanism, trained from scratch on the BabyLM 2026 Multilingual track corpus.
Data
100M tokens of equally-mixed English/Dutch/Chinese
~75M English-equivalent words (after Byte Premium adjustment)
Tokenizer: Regex-Guided BBPE (16K vocab)
Architecture
8 layers, hidden dim=512, 8 attention heads
Max sequence length: 512
~78M parameters
Training
10 epochs
Optimizer: AdamW (peak lr=3e-4, cosine decay, warmup 1000 steps)
Mixed precision: bfloat16
Final loss: ~2.3-2.5