-
📝 Rule-based Reward
✔️ Checks correctness of function call name and arguments.
➕ Partial credit for matching subsets of arguments.
-
🔒 Self-Certainty Reward
⚡ Encourages confident predictions.
-
🔧 Tool-Call Reward
✅ Validates structural correctness.
-
Why it lower than technical report?
There have a limit of hardware so have to reduce some max tokens when evaluation for both 2 models
-
Fair evaluate ?
I use the same configuration for all the models I review for larger or with a same size model.
I would be happy to receive a contribution to this model and get feedback about performance, quality of model
1@misc{qwen3-4b-i-1509,
2 title = {Qwen3-4B-I-1509: Fine-tuned Qwen3-4B-Instruct with GRPO for Tool-Use and Function Calling},
3 author = {Beyoru},
4 year = {2025},
5 howpublished = {\url{https://huggingface.co/beyoru/Qwen3-4B-I-1509}}
6}
7