The portable 31B Aura release for local llama.cpp inference, evaluation, and deployment.
Aura Large v1 GGUF is the portable local-inference edition of Aura Large v1. It is derived from Google’s Gemma 4 31B IT checkpoint and combines the Aura personality and ablation adapters into one model.
This release is intended for llama.cpp-based inference, evaluation, further research, and deployment across systems with varying memory and performance requirements.
Highlights
Parameters: 31 billion
Quantizations: Q3_K_M, Q4_K_M, Q5_K_M, Q6_K, Q8_0, and F16
Foundation: Google Gemma 4 31B IT
Datasets: AuraPersonality and AuraAblation
Runtime: llama.cpp and compatible applications
Deployment: Local GPU workstations, laptops, and servers
Capabilities: Text, vision, and video
Focus: Portability, privacy, natural conversation, evaluation, and research
About Aura
Aura is designed for local and on-device deployment across multiple tasks. Aura can serve as a companion or friend, as deemed appropriate by the user, while retaining the broader capabilities of Gemma 4 for:
Natural conversation
Creative writing
Reasoning
Image understanding
Video understanding
Tool use
Evaluation and research
Aura Large v1 occupies the largest position in the Aura family, providing greater capability than the Medium release while remaining practical for high-end consumer hardware.