🧠 dpp-gpt V2.1 Base (90m)
(🇺🇸 English / 🇷🇺 Русский)
⚠️
Note: This is a base foundation model. It has only undergone pre-training and has not been fine-tuned for general chat. For the instruction-following / chat version, please download dpp-gpt-V2.1-Flash-90m.
This is a microscopic foundation language model with 93M parameters, trained entirely from scratch. Despite its tiny size, it possesses native mathematical and translation capabilities embedded directly into its pre-training weights.
⚙️ Model Details
- Parameters: 93M
- Layers / Hidden Size / Heads: 11 / 768 / 12
- Context Length: 4096 tokens
- Vocabulary Size: 32,768
- Type: Base (Pre-trained foundation model)
- Format: GGUF / PyTorch (.pth)
- License: Apache 2.0
📊 Pre-training Data
- Dataset Size: ~11.26 Billion tokens (~121.1 tokens/parameter).
- Data Distribution:
- 💻 Code & Programming (Python, etc.): ~30%
- 🌍 Wikipedia (RU/EN/FR): ~20%
- 📚 CulturaX Corpus: ~15%
- 🧠 Cosmopedia (Synthetic textbook style): ~10%
- 🧮 Math, Logic, Translation datasets: ~25%
💡 Prompt Guide & Real Examples
As a raw base model, it is highly sensitive to syntax. Mathematical and textual formatting was baked into the pre-training weights. Use strictly the following structures to get accurate results:
🌍 Basic Translation
⚠️ Tip: This model is highly sensitive to the user's prompt formatting. To get a stable and correct translation, you must start your prompt with a capital letter and experiment with punctuation (sometimes adding or removing a period at the end changes the output).
1<|im_start|>user
2Переведи на английский: Я тебя люблю.<|im_end|>
3<|im_start|>assistant
4I love you.
🧮 Math Calculation (Chain-of-Thought)
To activate step-by-step math solving, use the [THINK] token:
1<|im_start|>user
2[THINK] 1523 - 659 + 234<|im_end|>
3<|im_start|>assistant
(The model will continue by breaking down the calculation into Thousands, Hundreds, Tens, and Units).
🇷🇺 Описание на русском
⚠️
Внимание: Это базовая (foundation) модель. Она прошла только этап pre-training и не обучалась свободному ведению диалога. Если вам нужна готовая чат-версия, скачайте dpp-gpt-V2.1-Flash-90m.
Это базовая компактная языковая модель на 93М параметров, обученная полностью с нуля. Даже в сыром виде без файнтюнинга модель демонстрирует базовые навыки пошагового счета и перевода, зашитые напрямую в веса предобучения.
⚙️ Детали модели
- Параметры: 93M
- Слои / Размерность / Головы: 11 / 768 / 12
- Контекст: 4096 токенов
- Словарь: 32,768 токенов
- Формат: GGUF / PyTorch (.pth)
- Лицензия: Apache 2.0
📊 Обучающие данные (Претрейн)
- Объем датасета: ~11.26 млрд токенов (~121.1 токена на параметр).
- Состав датасета:
- 💻 Код и программирование (в т.ч. Python): ~30%
- 🌍 Википедия (RU/EN/FR): ~20%
- 📚 Корпус CulturaX: ~15%
- 🧠 Cosmopedia (синтетические учебные тексты): ~10%
- 🧮 Математика, логика и параллельные корпуса для перевода: ~25%
💡 Важное руководство по промптам (Prompt Guide)
Так как это сырая базовая модель, она крайне чувствительна к синтаксису запроса. Если формат нарушен — логика генерации сломается. Разметка зашивалась в претрейн через формат ChatML, поэтому используйте строго следующие конструкции:
🌍 Базовый перевод
⚠️ Совет: Модель крайне чувствительна к формату вашего запроса. Чтобы получить стабильный и правильный перевод, пользователю необходимо начинать фразу с заглавной буквы, а также иногда экспериментировать с пунктуацией (наличие или отсутствие точки в конце запроса может поменять результат).
1<|im_start|>user
2Translate to russian: I love you.<|im_end|>
3<|im_start|>assistant
4Я люблю тебя.
🧮 Математический расчет (Chain-of-Thought)
Для активации пошагового решения используйте тег [THINK]:
1<|im_start|>user
2[THINK] 1523 - 659 + 234<|im_end|>
3<|im_start|>assistant
(Модель сама продолжит текст, расписав вычитание и сложение по разрядам).
Hotfix
The GGUF files in this repository were re-uploaded as a hotfix.
An LM Studio update changed the way special tokens are handled, which made the
previously published GGUF files generate broken output. The conversion has been
fixed and the quantizations here were rebuilt from the corrected model. The
weights are unchanged — only the token metadata inside the GGUF files.
If you downloaded a GGUF from this repository before this commit, please
download it again.