High-quality 5.26 bpw GGUF quantization of
ressl/MiniMax-M3-uncensored,
optimized for character-heavy fantasy roleplay and long-context continuity.
That checkpoint is an uncensored/abliterated derivative of
MiniMaxAI/MiniMax-M3.
[!CAUTION]
Risk level: high. This model is genuinely uncensored. The source reports
zero hard refusals on its harmful-prompt sample. Quantization does not restore
safety alignment. The model can generate dangerous, illegal, abusive,
sexually explicit, manipulative, or otherwise harmful content, including
operational cyber or violence-related instructions. Do not expose it to
untrusted users without strong access controls, monitoring, policy
enforcement, and output safeguards. You are responsible for lawful and safe
use.
What this is
This is a quantized derivative, not a new finetune. No gradient training was
performed. A roleplay-domain importance matrix was used to choose quantization
scales while the final GGUF was produced directly from the ReSSl BF16 weights.
Format: 18-shard, text-only GGUF
Size: 280,052,357,504 bytes (260.819 GiB)
Effective precision: 5.26 bits per weight
Architecture: MiniMax-M3, 428B total / approximately 23B active MoE
Context metadata: 1,048,576 tokens
Tested context: 65,536 tokens
Native MiniMax-M3 chat template retained
Vision tower and MTP head omitted
Start with MiniMax-M3-uncensored-RP-LC-Q4HQ-00001-of-00018.gguf;
llama.cpp discovers the other shards automatically.
Quantization recipe
This is deliberately more conservative than a conventional all-Q4 build:
Tensor group
Type
Count
Routed expert gate/up
Q4_K
114
Attention, routed-down, shared/dense expert paths
Q6_K
477
MSA indexer Q/K projections
F32
114
Token embedding and output
Q8_0
2
Routers and normalization tensors
F32
retained
Residual-writing down projections, all attention projections, shared experts,
and the first three dense MLP blocks were kept at Q6_K. MiniMax-M3's sparse
attention indexers were retained at F32 because they directly determine which
long-context blocks remain visible.
Roleplay calibration
The quantization importance matrix used 655,360 tokens from a deterministic
fantasy/creative roleplay corpus:
Component
Windows
Tokens
Share
Short-window matrix
128 x 4,096
524,288
80%
Long-window supplement
8 x 16,384
131,072
20%
The source pool mixed character dialogue, multi-character roleplay, creative
fiction, and narrative worldbuilding from
Timersofc/creative-writing-reap-calibration
and agentlans/combined-roleplay.
The combined matrix contained 762 tensor entries with 98.44% minimum routed-
expert coverage.
This calibration changes quantization error allocation only. It does not add
new roleplay training or alter the source model's behavior through finetuning.
Evaluation
Held-out perplexity
Four document-disjoint 4,096-token fantasy-roleplay chunks:
Model
PPL
Relative to Q8
Q8 reference
3.7610 +/- 0.09740
baseline
Earlier 4K-only RP-Q4HQ
3.7819 +/- 0.09923
+0.56%
This RP-LC-Q4HQ
3.7669 +/- 0.09864
+0.16%
Long-context continuity
A deterministic 61,500-token user prompt became 61,709 tokens after the
native chat template. Six trusted campaign records were placed from token 124
through token 53,321 among untrusted fantasy-roleplay distractors.
Exact canon recall: 14/14
Continuation length: 541 words (requested range: 450-650)
Canon facts used naturally in the scene: 8 (requested minimum: 6)
Distinct NPC voices: pass
Player-character agency preserved: pass
Concrete action opening at the ending: pass
On a 512 GiB Apple M3 Ultra with a 65,536-token context:
Prompt processing: 189.5 tokens/s
Generation: 16.1 tokens/s
Peak process RSS: 271.05 GiB
Process swaps: 0
These are targeted quantization checks, not a comprehensive capability or
safety evaluation. The 1M context declared by the architecture was not tested.
Runtime
The model requires MiniMax-M3's trained MSA sparse-attention implementation. It
was built and validated with llama.cpp build 10018 at commit
e99545c1c41ba42b7986831c0de2983498dc3c5b,
from the MiniMax-M3 MSA support work.
Use that revision or a later llama.cpp version with equivalent MiniMax-M3 MSA
support. Substituting dense attention is not equivalent beyond the dense
prefix.
bash
1huggingface-cli download m9e/MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF \2 --local-dir MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF
34llama-cli \5 -m MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF/MiniMax-M3-uncensored-RP-LC-Q4HQ-00001-of-00018.gguf \6 -ngl all -fa on -fit off -c 65536 -cnv \7 -rea off --reasoning-budget 0\8 --temp 1.0 --top-p 0.95 --min-p 0.05
The complete model occupies about 261 GiB before runtime buffers and context.
The tested 65K configuration peaked near 271 GiB. Plan hardware capacity with
additional headroom; smaller context allocations reduce runtime overhead.
Intended uses
Private or access-controlled creative writing and fantasy roleplay
Long-running campaign continuity experiments
Controlled research into uncensored model behavior
Lawful, authorized security research and red-team analysis with safeguards
Safety, limitations, and out-of-scope use
The source checkpoint intentionally suppresses refusal behavior. Treat model
output as untrusted data and never as authorization to act.
Do not use it to facilitate crime, malware deployment, unauthorized access,
violence, self-harm, exploitation, privacy abuse, harassment, or deception.
Roleplay output may include graphic violence, sexual material, coercion, or
other disturbing themes. Use explicit consent and age/content controls. Never
use it to sexualize or exploit minors.
Do not rely on it for medical, legal, financial, emergency, or other
high-stakes decisions.
It may hallucinate facts, lose continuity, imitate copyrighted styles, or
reproduce biases from its source models and calibration text.
This GGUF is not a security boundary. Applications should add authentication,
rate limits, logging, abuse monitoring, prompt/output filtering, and human
review appropriate to their risk.
The MiniMax Community License contains binding prohibited-use and commercial-
use conditions. Read LICENSE before use or redistribution.
Lineage and credits
Original model:MiniMaxAI/MiniMax-M3,
by MiniMaxAI. Source revision used by the ReSSl derivative:
50942730318c7943fe83db7ec8e9f9177ecb1cf8.
Uncensored BF16 source:ressl/MiniMax-M3-uncensored,
uncensoring and validation by Robert Ressl. Revision:
315b596663fbe37fce3880d7c0468ceb47cd2da5.
GGUF conversion, calibration, quantization, and evaluation:m9e.
Thanks to the authors and maintainers of the calibration datasets named above.
License
This derivative inherits the MiniMax Community License from the original
model. The license is included in this repository and is also available in the
upstream MiniMax-M3 repository.
It includes attribution, commercial-use, and prohibited-use requirements. This
model is provided as-is, without warranty.
Integrity
SHA256SUMS contains a digest for every GGUF shard. Verify after download: