Qwen3.5-2B is a 2-billion-parameter, dense Transformer decoder-only language model developed by the Qwen team at Alibaba Cloud, featuring a
3:1 DeltaNet linear-attention to full-attention layer ratio for efficient long-context processing with
256K native context length. Pre-trained on diverse public corpora and aligned via SFT + DPO. This distribution is provided by
Aria Compute as an
aria-quant-bundle — a uniform 8-bit per-group quantized package using
Hadamard pre-processing + uniform quantization with per-group codebooks (g=32). It delivers the
near-lossless option in the Aria Compute lineup — at ~1.9× smaller than BF16 (~4 GB → ~2.1 GB) with negligible quality degradation (logprob delta +0.00685 on method reference). Optimized for
CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the
Aria Engine runtime. No GPU or cloud connection is required.
Authenticated dashboard users can download the bundle via:
https://ariacompute.com/dashboard/models
Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend: