The mobile-optimized Aura release for supported on-device and edge hardware.
Aura v1.1 LiteRT-LM is built from Google’s Gemma 4 E4B mobile-QAT checkpoint with the canonical combined Aura adapter baked into the model as floating-point residual branches.
The v1.1 release has seen significant improvements to the personality adapter, merged and applied with the same ablation adapter as v1.
This edition is intended for efficient multimodal deployment through LiteRT-LM on supported mobile and edge devices.
Highlights
Format: LiteRT-LM
Foundation: Gemma 4 E4B mobile-QAT
Adapter: Canonical Aura v1 composite, baked into the graph
Runtime: LiteRT-LM
Deployment: Supported mobile and edge hardware
Focus: Efficient, private, on-device inference
About Aura
Aura is designed for on-device deployment across multiple tasks. Aura can be a companion or friend, as deemed necessary by the user, while also supporting:
Natural conversation
Creative writing
Reasoning
Image understanding
Multimodal interaction
Private and offline workflows
Format Notes
The mobile-QAT base tensors retain their optimized deployment representation. Aura is applied through separate floating-point residual branches rather than being merged directly into the packed base weights.
This approach preserves the intended mobile-QAT conversion path while retaining the behavior of the canonical combined Aura adapter.