This model is a fine-tuned version of
unsloth/Qwen3.5-9B, optimized for agent-based reasoning tasks. It was trained using the
Unsloth framework to achieve faster training speeds and memory efficiency.
This model was trained
2x faster using
Unsloth combined with Hugging Face's TRL library. Unsloth allows for efficient fine-tuning of Large Language Models (LLMs) with significantly reduced VRAM usage and increased throughput.