Views
No views yet
tacita-desktop (Rust runtime). For mobile (flutter_gemma), use the LiteRT‑LM build: Gemma‑4‑Tacita‑E4B‑litert‑lm.enable_thinking honored natively; emits nothing when thinking is off (fixes the stock‑Gemma llama.cpp #21338 class of bugs).{queries:[{q,lang}]}), grounds answers, retries off‑topic results.tacita.* keys (gguf namespace) so the Tacita desktop runtime reads them once at load and bypasses the matching director call. A stock llama.cpp runtime ignores them and runs the model normally.tacita.model = gemma-4-tacita
tacita.variant = E4B
tacita.tier = desktop-pro
tacita.capabilities = [preamble_inline, thinking_native, search_plan_inline, ...]| File | Bits | Size (approx) | Use |
|---|---|---|---|
*-Q4_K_M.gguf | 4‑bit | ~2.6 GB | default desktop |
*-Q8_0.gguf | 8‑bit | ~4.3 GB | quality‑first |
*-F16.gguf | 16‑bit | ~8 GB | reference |
llama.cpp (b‑recent) or any GGUF runtime. eos_token is <end_of_turn> (Gemma‑4 turn format).