Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
smol_llama-220M-GQA-bpw3.5 – AI Model by blockblockblock | AlphaNeural AI
You can deploy this model and start earning money today!
blockblockblock
/
smol_llama-220M-GQA-bpw3.5
like
0
transformers
llama
text-generation
smol_llama
llama2
en
JeanKaddour/minipile
pszemraj/simple_wikipedia_LM
mattymchen/refinedweb-3m
BEE-spoke-data/knowledge-inoc-concat-v1
apache-2.0
model-index
autotrain_compatible
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
smol_llama: 220M GQA
model card WIP, more details to come
A small 220M param (total) decoder model. This is the first version of the model.
1024 hidden size, 10 layers
GQA (32 heads, 8 key-value), context length 2048
train-from-scratch on one GPU :)
Links
Here
are some fine-tunes we did, but there are many more possibilities out there!
instruct
openhermes -
link
open-instruct -
link
code
python (pypi) -
link
zephyr DPO tune
SFT -
link
full DPO -
link
Open LLM Leaderboard Evaluation Results
Detailed results can be found
here
Metric
Value
Avg.
29.44
AI2 Reasoning Challenge (25-Shot)
24.83
HellaSwag (10-Shot)
29.76
MMLU (5-Shot)
25.85
TruthfulQA (0-shot)
44.55
Winogrande (5-shot)
50.99
GSM8k (5-shot)
0.68