Qwen3.5-0.8B is a 0.8-billion-parameter, dense Transformer decoder-only language model developed by the Qwen team at Alibaba Cloud, featuring a
3:1 DeltaNet linear-attention to full-attention layer ratio for efficient long-context processing. Pre-trained on diverse public corpora and aligned via SFT + DPO. This distribution is provided by
Aria Compute as an
aria-quant-bundle — a uniform 4-bit quantized package using
Hadamard rotation + Lloyd-Max codebook quantization with per-group codebooks (group size 32). It delivers the
smallest practical bundle — at ~3.6× smaller than BF16 (~1.6 GB → ~450 MB) — for maximum compression when disk and memory footprint are critical. Optimized for
CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the
Aria Engine runtime. No GPU or cloud connection is required.
Authenticated dashboard users can download the bundle via:
https://ariacompute.com/dashboard/models
Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend: