Poro 2 8B Math Reasoning RL Preview is a specialized model focused on mathematical reasoning and problem-solving. This preview model was created through reinforcement learning (RL) on top of the Math Reasoning SFT checkpoint. This model excels at mathematical reasoning tasks but is not optimized for general conversational use or other domains.
The model produces reasoning traces in both English and Finnish, adapting the language of its reasoning based on the language of the input prompt. The RL training used separate environments for English and Finnish math problems, a novel approach that maintains strong bilingual reasoning performance.
The Poro 2 Long model family extends the Poro 2 models with longer context support and checkpoints trained especially on math reasoning. There are four checkpoints released: a base model, an instruction-tuned model, a math reasoning SFT checkpoint, and the final math reasoning RL checkpoint.
Poro 2 8B Math Reasoning RL Preview is based on the Llama 3.1 8B architecture and has been fine-tuned for mathematical reasoning through supervised fine-tuning followed by reinforcement learning.
Poro 2 8B Math Reasoning RL Preview demonstrates strong mathematical reasoning capabilities in both English and Finnish, showing substantial improvements over both the SFT checkpoint and the original Poro 2 8B Instruct model. On AIME 2025, the model achieves performance comparable to models 6x its size.
Mathematical reasoning in both English and Finnish
Note: This model is optimized specifically for mathematical reasoning and is not suitable for general conversational AI or other domains.
Ethical Considerations and Limitations
Poro 2 8B Math Reasoning RL Preview is a specialized model optimized for mathematical reasoning in English and Finnish. As a preview release, this model has specific limitations:
Key limitations:
Optimized exclusively for mathematical reasoning; not suitable for general conversation, coding, creative writing, or other domains
This is an early preview release
Reasoning traces are in English or Finnish depending on prompt language; limited proficiency in other languages
May occasionally generate biased, inappropriate, or factually incorrect content
Safety Considerations:
Users should verify important factual claims independently
The model should not be used for medical, legal, or financial advice without human oversight
Responses should be reviewed for appropriateness in sensitive contexts
License
Built with Llama
Poro 2 8B Math Reasoning RL Preview is released under the Llama 3.1 Community License. Please review the license terms before use.
We thank CSC - IT Center for Science, Finland for providing access to the LUMI supercomputer, and TensorWave for providing access to AMD MI325X GPU clusters. This work was supported by the High Performance Language Technologies (HPLT) project and conducted in collaboration with TurkuNLP from the University of Turku. This project has received funding from the European Union's Horizon Europe research and innovation programme under grant agreement No 101070350.