This is a
W4A16 (4-bit weight, 16-bit activation) quantized version of
Qwen/Qwen3.5-0.8B, produced using
AutoRound — Intel's sign gradient descent based quantization method designed for production-grade accuracy retention.
This model is compatible with
transformers and backends that support AutoRound format weights (e.g., vLLM, SGLang). For full model details, architecture, and capabilities, refer to the
base model page.
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.