CURE-MED-32B is a 32 billion parameter large language model specialized for multilingual medical reasoning, fine-tuned from Qwen/Qwen2.5-32B using a
curriculum-informed reinforcement learning framework to enhance logical correctness and language stability in healthcare applications.
CURE-MED-32B is part of the CURE-MED family of models, designed to address the challenges of multilingual medical reasoning in large language models (LLMs).
Built on the Qwen/Qwen2.5-32B-Instruct model, it incorporates a curriculum-informed reinforcement learning approach that integrates code-switching-aware supervised fine-tuning (SFT)
and Group Relative Policy Optimization (GRPO) to improve performance on open-ended medical queries across 13 languages, including underrepresented ones such as Amharic, Yoruba, and Swahili.
The model is trained and evaluated using CUREMED-BENCH, a high-quality multilingual open-ended medical reasoning benchmark with single verifiable answers.
This is the model card of a 🤗 transformers model that has been pushed on the Hub.
1@article{onyame2026cure,
2 title={CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning},
3 author={Onyame, Eric and Ghosh, Akash and Baidya, Subhadip and Saha, Sriparna and Chen, Xiuying and Agarwal, Chirag},
4 journal={arXiv preprint arXiv:2601.13262},
5 year={2026}
6}
7
8