Gemopus 4 26B A4B IT Local Abliterated SOTA Internal R7 Selected T34 Transfer
This checkpoint was produced with model-forge from Jackrong/Gemopus-4-26B-A4B-it. It is not a new fine-tune. It computes fresh refusal directions on the already fine-tuned Gemopus model, then applies the selected base Gemma Heretic t34 per-layer parameters.
Recipe
Generated with
model-forge, a model-agnostic post-training pipeline for fine-tuning, refusal ablation, evaluation, and publishing.
Repository recipe: configs/abliteration/gemma4_26b_a4b_ft_local_abli.yaml
Transferred base recipe: configs/abliteration/gemma4_26b_a4b_local_abli.yaml, [Trial 34] Refusals: 1/27, KL divergence: 0.0183.
Key settings: Heretic direct-parameter transfer, source model Jackrong/Gemopus-4-26B-A4B-it, fresh refusal directions from model-forge internal prompts on the FT model, per-layer direction scope, full row normalization, orthogonalized refusal direction.
Evaluation
| Bucket | Metric | Score |
|---|
| refusal_calibration_unsafe | ablation_refusal_suppression_rate | 1.0 |
| refusal_paired_boundary | ablation_refusal_suppression_rate | 1.0 |
| unsafe_overcompliance | ablation_refusal_suppression_rate | 0.6667 |
| capability_preservation_challenge | normal_use_regression_pass_rate | 0.7812 |
| normal_use_regression | normal_use_regression_pass_rate | 1.0 |
| refusal_paired_boundary | benign_answer_quality_rate | 0.50 |
The FT source model scored 0.7812 on capability_preservation_challenge, 1.0 on normal_use_regression, and 0.50 on paired benign answer quality in the same local eval suite. This ablation preserved those FT performance levels while reducing refusals on the targeted unsafe gates.
Intended Use
This model is intended for controlled ablation research and evaluation of post-training/refusal-removal recipes. It may comply with unsafe requests more often than the source instruction-tuned model.
Safety, Provenance, and Use Guidance
This is a refusal-ablated research derivative. The modification intentionally changes refusal behavior and can reduce safeguards present in the source model. It should not be treated as safety-aligned merely because the upstream model included safety tuning.
Use it only in controlled research or evaluation environments with independent content controls, logging, access restrictions, and task-specific safety testing. Do not deploy it as an unreviewed public assistant or in high-impact domains.
The source model and transferred recipe are identified above. The repository records the Model Forge procedure and diagnostics, but the exact upstream source commit was not pinned in the original public card. Preserve the current artifact and record that revision before any future rebuild or derivative release.