Views
No views yet
google/gemma-4-12b-it,
built for vLLM on NVIDIA V100 (SM70) via the
1Cat-vLLM fork.google/gemma-4-12B-it-qat-w4a16-ct), which are QAT-trained and likely
higher quality. This RTN checkpoint remains useful as a smoothing-free
baseline.1from vllm import LLM
2llm = LLM(model="JelleFoks/gemma-4-12b-it-AWQ-rtn", dtype="float16",
3 trust_remote_code=True)sm70-v100-rebase-new: TG B=1 37.8 / B=8 186.7 tok/s,
PP ctx=512 1698 tok/s (V100-32GB, 1380 MHz).