This repo contains a GGUF Q4_K_M quantized version of the merged v5 model for smaller downloads and cross‑platform CPU/GPU inference via llama.cpp / llama-cpp-python.
Files:
latentscore-gemma3-270m-v5-q4_k_m.gguf
Notes:
This is intended for local inference with llama.cpp‑compatible runtimes.
Chat formatting should follow the model’s chat template (Gemma 3 / Unsloth style).