A curated dataset of reasoning traces for training local AI orchestrator agents.
Designed for SFT and RL training of Nemotron 3 Nano Omni to be the best local
Hermes Agent model.
Total SFT rows: 28,000
Total RL prompts: 28,000
Format: ShareGPT (conversations column)
Target model: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
Training framework: Unsloth Studio