This is an uncensored version of
google/gemma-3-27b-it created with a new abliteration technique.
See
this article to know more about abliteration.
I was playing with model weights and noticed that Gemma 3 was much more resilient to abliteration than other models like Qwen 2.5.
I experimented with a few recipes to remove refusals while preserving most of the model capabilities.
Note that this is fairly experimental, so it might not turn out as well as expected.
In the original technique, a refusal direction is computed by comparing the residual streams between target (harmful) and baseline (harmless) samples.
Here, the model was abliterated by computing a refusal direction based on hidden states (inspired by
Sumandora's repo) for each layer, independently.
This is combined with a refusal weight of 1.5 to upscale the importance of this refusal direction in each layer.
This created a very high acceptance rate (>90%) and still produced coherent outputs.