GRPO-reinforced LoRA adapter for Nemotron — the RL stage that aligns tool-use with ontology constraints.
Sangue e Grafi banner
Model Description
This is a LoRA (PEFT) adapter trained via Group Relative Policy Optimization (GRPO) on top of the Nemotron SFT adapter (already merged into nvidia/Nemotron-Mini-4B-Instruct). It is the second training stage for the Nemotron-based pipeline in the Sangue e Grafi project.