Views
No views yet
Qwen/Qwen3.5-4B, produced with orthex — an implementation of the technique from Arditi et al., "Refusal in Language Models Is Mediated by a Single Direction" (NeurIPS 2024).qwen3_5resid_preembed_tokens, and every layer's attn_out and mlp_out — orthogonalized in place in the weights (not a runtime hook; this checkpoint behaves this way standalone, with no orthex dependency at inference time)| metric | pre | post | delta |
|---|---|---|---|
| refusal rate | 0.969 | 0.000 | -0.969 |
| perplexity | 11.211 | 11.248 | 0.037 |
evaluation_report.json in this repo for the full per-prompt breakdown (refusal_samples) and the ranked candidate list considered during selection (selection_report).Qwen/Qwen3.5-4B's license allows.Qwen/Qwen3.5-4B's license; set this field explicitly before publishing.