This model is a continual pretraining (CPT) checkpoint built by further pretraining Qwen2.5 3B on the Uyghur portion of the
MC^2 Corpus.
The model is intended to improve Uyghur language modeling and to support research on low-resource language adaptation.
This checkpoint is released primarily for research purposes. Researchers are welcome to use this CPT checkpoint as a base model for future work, particularly in model merging and logit fusion.
1@article{zhang2026efficient,
2 title={Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion},
3 author={Zhang, Chen and Lin, Jiuheng and Liao, Zhiyuan and Feng, Yansong},
4 journal={arXiv preprint arXiv:2604.18106},
5 year={2026}
6}