We release GreenMind-Medium-14B-R1, a medium-sized Vietnamese language model capable of effectively addressing questions that require intermediate-level reasoning, such as general knowledge, mathematics, natural science and social science topics. By leveraging the Group Relative Policy Optimization strategy for fine-tuning, we guide the model to generate logically coherent responses.
Context Length: Full 131,072 tokens and generation 8192 tokens
Language: Vietnamese
Quickstart
Here provides a code snippet with apply_chat_template to show you how to load the tokenizer and model and how to generate contents.
python
1from transformers import AutoModelForCausalLM, AutoTokenizer
23model_name ="GreenNode/GreenMind-Medium-14B-R1"45model = AutoModelForCausalLM.from_pretrained(6 model_name,7 torch_dtype="auto",8 device_map="auto"9)1011tokenizer = AutoTokenizer.from_pretrained(12 model_name,13 revision='main',14 trust_remote_code=False,15)16prompt =r"""Vừa gà vừa chó
17Bó lại cho tròn
18Ba mươi sáu con
19Một trăm chân chẵn
20Hỏi có bao nhiêu con gà, bao nhiêu con chó?"""2122messages =[23{24"role":"system",25"content":"Bạn là một trợ lý ảo hữu ích trong việc trả lời câu hỏi. Hãy suy luận từng bước, và đưa ra đáp án trong thẻ <answer> </answer>."26},27{28"role":"user",29"content":f"{prompt} Hãy suy luận từng bước trong thẻ <think> </think>. Và trả về đáp án trong thẻ <answer> </answer>."30},31{32"role":"assistant",33"content":"Hãy để tôi giải quyết từng bước.\n<think>"34}35]3637text = tokenizer.apply_chat_template(38 messages,39 tokenize=False,40 continue_final_message=True)4142model_inputs = tokenizer([text], return_tensors="pt").to(model.device)4344generated_ids = model.generate(45**model_inputs,46 max_new_tokens=102447)4849generated_ids =[50output_ids[len(input_ids):]for input_ids, output_ids inzip(model_inputs.input_ids, generated_ids)51]5253response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]54print(response)55# Đầu tiên, chúng ta cần thiết lập hai phương trình dựa trên thông tin đề bài:56# 1. Tổng số con gà và chó là 36: x + y = 3657# 2. Tổng số chân là 100: 2x + 4y = 10058# Trong đó, x là số con gà và y là số con chó.59# Tiếp theo, chúng ta giải hệ phương trình này:60# Từ phương trình thứ nhất, ta có: x = 36 - y61# Thay vào phương trình thứ hai: 2(36 - y) + 4y = 10062# => 72 - 2y + 4y = 10063# => 2y = 2864# => y = 14 (số con chó)65# Thay y = 14 vào phương trình x + y = 36:66# => x = 36 - 14 = 22 (số con gà)67# Vậy, có 22 con gà và 14 con chó.68# </think>69# <answer>Có 22 con gà và 14 con chó.</answer>
Evaluation
Table 1. SeaExam Dataset. GreenMind-Medium-14B-R1 compared to base model and some models with larger size.
Model
SeaExam-ID
SeaExam-TH
SeaExam-VI
Avg
Meta-Llama-3.1-70B-Instruct
65.8
70.6
72.6
69.7
gemma3-27b-it
64.4
67.5
73.1
68.4
Qwen2.5-14B-Instruct
67.6
68.8
73.1
69.8
GreenMind-Medium-14B-R1
74.36
69.75
74.44
72.79
Table 2. VLSP 2023 Challenge: The performance of our model outperforms most SOTA models.
Model
ComprehensionQA-vi ↑
Exams-vi ↑
LAMBADA-vi ↓
WikiQA-vi ↑
MMLU-vi ↑
cpt-smartbot-13b
0.6633
0.3473
21.9864
0.4455
0.414
ura-llama-13b
0.6556
0.342
17.5614
0.438
0.3973
greennode-7b (prior work)
0.6122
0.2892
189.7782
0.3335
0.387
greennode-14b (prior work)
0.6711
0.3672
29.5967
0.468
0.5281
GreenMind-Medium-14B-R1 (Ours)
0.8689
0.7796
10.7609
0.7915
0.7124
Table 3. VMLU Dataset. The performance compared to fine-tuned models.
This repository and the model weights are licensed under the MIT License.
Citation
If you find our work helpful, feel free to give us a cite.
@misc{tung2025greenmindnextgenerationvietnameselarge,
title={GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning},
author={Luu Quy Tung and Hoang Quoc Viet and Pham Bao Loc and Vo Trong Thu},
year={2025},
eprint={2504.16832},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2504.16832},
}