Gemma-3n-E2B-IT is a ~2-billion-parameter instruction-tuned multimodal language model developed by Google, part of the Gemma 3n family. Its text backbone features
hybrid attention (24 of 30 layers sliding-window linear attention with activation sparsity + 6 Laurel low-rank full-attention layers),
GeGLU activation,
Grouped Query Attention (GQA),
per-layer input projections, and
32K native context length. Pre-trained on diverse web-scale corpora and aligned via instruction tuning + RLHF. This distribution is provided by
Aria Compute as an
aria-quant-bundle — a quantized package using
Hadamard pre-processing + per-channel quantization. Optimized for
CPU-only, on-device inference on mobile phones, edge devices, and single-board computers via the
Aria Engine runtime. No GPU or cloud connection is required.
Authenticated dashboard users can download the bundle via:
https://ariacompute.com/dashboard/models
Users (both direct and downstream) should be made aware of the above risks, biases, limitations, and constraints of the model. We recommend: