AWQ W4A16 (group_size 128, symmetric) quantization of
trohrbaugh/gemma-4-31b-it-heretic-ara
— a Heretic-ARA abliterated derivative of
google/gemma-4-31b-it.
1vllm serve alonsoko/gemma-4-31b-it-abliterated-heretic-AWQ-W4A16 \
2 --trust-remote-code \
3 --tensor-parallel-size 1 \
4 --max-model-len 32768
This model is shared purely for academic and technical exploration of model internals.
Gemma 4 31B is a dense multimodal model (text + image input, text output) with
a 256K context window, native thinking-mode support, function calling, and
strong performance on reasoning, coding, and vision benchmarks. See the
base model card for architectural
details, benchmark results, training data, and Google's responsible-use guidance.