This is
Qwen/Qwen3-4B-Thinking-2507 quantized with
LLM Compressor with NVFP4. The model has been created, tested, and evaluated by The Kaitchup.
The model is compatible with vLLM v0.11. Tested with an RTX 5090.
Subscribe to
The Kaitchup. This helps me a lot to continue quantizing and evaluating models for free. Or if you prefer to give some GPU hours, "
buy me a kofi"