Views
No views yet
THUDM/glm-4-9b-chat-hf, produced with Solutus
(measurement-first LLM abliteration). Private research artifact — outputs are the base model's, minus
the refusal behavior; use responsibly.directional (single refusal direction), whitened-SVD extraction, byte-exact (batch_size=1):solutus abliterate THUDM/glm-4-9b-chat-hf --technique directional \
--dataset advbench,harmbench,multijail_zh,sorrybench \
-o extraction=whitened_svd -o n_directions=1 --max-new-tokens 512n_directions=4
over-ablated (WikiText ΔPPL +30% vs +9.8% smoke), so n_directions=1 is the capability-preserving recipe.| Axis | Result |
|---|---|
| refusal — advbench / harmbench | 15.6% / 6.2% |
| refusal — MultiJail zh / ar / sw | 0% / 3.1% / 0% |
| over-refusal — orbench_hard (benign) | 0% refusal (stays benign-compliant) |
| coherent-compliance | ~90–100% (see Swahili caveat) |
| capability — WikiText-2 ΔPPL | −0.4% (base 29.13 → 29.00 — no degradation) |
| capability — GSM8K / MMLU (n=100) | 59.0% / 67.0% |