A 48.5M parameter DeepSeek-V4 architecture model trained from scratch on a
multilingual blend of C4 data (108 languages). This is a research sandbox
model for studying the DeepSeek-V4 architecture at small scale — it is
not a production model and has no practical use.
Based on the
DeepSeek-V4 architecture,
consumed via HuggingFace
DeepseekV4ForCausalLM (transformers ≥ 5.14):
This model was trained from scratch (random initialization) — it does not
inherit weights from any pretrained model.
Uses the DeepSeek-V4-Flash-0731 tokenizer (vocab 129,280).
This model exists purely for architectural experimentation. At 48.5M
parameters, it cannot perform any useful tasks. Do not use it for anything.