Views
No views yet

[!NOTE] Qwen3.5-9B-DeepSeek-V4-Flash is an efficient reasoning model distilled using high-quality data from DeepSeek-V4.

[!IMPORTANT] This is an early controlled Q5_K_M comparison between Jackrong/Qwen3.5-9B-DeepSeek-V4-Flash and the official Qwen3.5-9B base model.This evaluation was completed by Kyle Hessling, who ran the same evaluation suite twice under the same local inference conditions: once on the DeepSeek-V4 distill model and once on the official Qwen3.5-9B base model.






temperature=0.7 to 1.0 (Use lower temperature for strict coding tasks, higher for creative reasoning)top_p=0.95A Note: My goal isn't just to detail a workflow, but to demystify LLM training. Beyond the social media hype, fine-tuning isn't an unattainable ritual—often, all you need is a Google account, a standard laptop, and relentless curiosity. All training and testing for this project were self-funded. If you find this model or guide helpful, a Star ⭐️ on GitHub would be the greatest encouragement. Thank you! 🙏
1@misc{jackrong_qwen35_9b_deepseek_v4_flash,
2 title = {Qwen3.5-9B-DeepSeek-V4-Flash},
3 author = {Jackrong},
4 year = {2026},
5 publisher = {Hugging Face}
6}