This is a
W4A16 (4-bit weight, 16-bit activation) GPTQ-format quantized version of
Qwen/Qwen3.5-0.8B, produced using
AutoRound — Intel's sign gradient descent based quantization method designed for production-grade accuracy retention.
This model is compatible with
transformers,
AutoGPTQ,
vLLM, and
SGLang — any backend supporting GPTQ-format weights works out of the box. For full model details, architecture, and capabilities, refer to the
base model page.
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.