vidya-2b
Recommended settings (GGUF / mobile apps)
This GGUF ships a plain ChatML chat template (Qwen3 thinking mode disabled). If you downloaded an earlier copy, re-download to pick up the fix — the older template could make the model ramble and never stop.
Recommended sampler settings:
temperature: 0.3
repeat_penalty: 1.1 (do not use 1.0 — small models can loop; avoid > 1.3)
repeat_last_n: 64
- stop token:
<|im_end|> (already the model's EOS)