KyrgyzBert2
This repository contains second-generation Kyrgyz-first encoder pretrained from scratch with a ModernBERT-style architecture.
Model
- Name:
KyrgyzBert2
- Family: KyrgyzBERT
- Architecture: ModernBERT-style encoder
- Parameter count: 98.0M
- Tokenizer: 32,768-token byte-level BPE
Notes
This is the simplified Hugging Face repository name for the model previously uploaded as a KyrgyzBERTv2 checkpoint.
Intended Use
This checkpoint is intended for base encoder pretraining, masked-language-model research, and downstream fine-tuning for Kyrgyz NLP. It is intended for research, reproducibility, and educational use by the Kyrgyz NLP community. It should not be used for high-stakes decisions or production deployment without separate validation for the target domain.
License and Usage
License metadata is set to other. The checkpoint is released for research and reproducibility. Downstream datasets and base models may have their own licenses or usage terms; users are responsible for following the corresponding dataset cards and upstream model licenses. The checkpoint is provided without warranty.