Qwen3 1.7B is a 1.7-billion-parameter, dense Transformer decoder-only language model developed by the Qwen team at Alibaba Cloud, pre-trained on diverse public corpora and aligned via supervised fine-tuning (SFT) and direct preference optimization (DPO). This distribution is provided by
Aria Compute as an
aria-quant-bundle — a uniform 8-bit quantized package using
Hadamard rotation + Lloyd-Max codebook quantization with per-group codebooks (group size 32). It delivers
near-lossless generation quality — on qwen3-0.6B, logprob delta is just +0.00685 above FP16 and exact prefix fraction ties the best quant recipe at 0.3854 — at ~1.8× smaller than FP16 (~3.4 GB → ~1.9 GB). Optimized for
CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the
Aria Engine runtime. No GPU or cloud connection is required.
Authenticated dashboard users can download the bundle via:
https://ariacompute.com/dashboard/models
Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend: