Original model: a Task Arithmetic merge of Mistral-Nemo-Instruct-2407, Lumimaid-v0.2-12B, and absolute-heresy — tuned for uncensored roleplay and creative writing. 12B params, Mistral architecture.
Good quality, best size/speed tradeoff. Recommended default.
F16 (23 GB) is the unquantized conversion — only needed if you want to re-quantize yourself.
Which file should I choose?
Pick the largest quant that fits in your RAM/VRAM with room to spare (leave ~1–2 GB for context).
8 GB RAM/VRAM → Q4_K_M
12 GB → Q5_K_M or Q6_K
16 GB+ → Q6_K or Q8_0
For roleplay/creative use, Q4_K_M is plenty — quality loss over Q8 is negligible in practice.
Prompt format
Mistral instruct format (primary):
<s>[INST] {prompt} [/INST]
Most front-ends (SillyTavern, LM Studio) apply this automatically when you select a Mistral / Mistral-Nemo template. The Lumimaid component was also trained with ChatML, so ChatML works too — but Mistral format is the safe default.
Recommended settings
Mistral-Nemo (this model's base) is temperature-sensitive — high temperature makes it incoherent. Keep it low:
Setting
Value
Note
Temperature
0.5 – 0.7
Do not exceed ~1.0. Lower = more coherent, higher = more creative.
Min-P
0.05
Primary truncation. Leave Top-P/Top-K off if using Min-P.
Repetition penalty
1.05 – 1.1
Or use DRY (0.8 / 1.75 / 2) if your front-end supports it.
Context
up to 16k reliable
Nemo's trained context is 128k but quality degrades well before that.
Start at temp 0.6 / min-p 0.05 and adjust from there.
Run it
llama.cpp:
llama-cli -m Mistral-Heretica-12B-GGUF-Q4_K_M.gguf --jinja -p "Write a short noir opening."
Download one file with the HF CLI (skip the rest):
ollama run hf.co/AnkitAI/Mistral-Heretica-12B-GGUF:Q4_K_M
Verified
Q4_K_M smoke-tested: loads correctly, follows instructions (valid JSON output), and produces coherent creative prose. ~5–6 tok/s generation on Apple Silicon (M-series).
License
The base model declares no explicit license. This merge inherits terms from its source models — notably Lumimaid-v0.2-12B, which is CC-BY-NC-4.0 (non-commercial). Treat this model as non-commercial unless you verify otherwise with the original authors. These are quantizations only; all model rights belong to the original creators.