Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen3.5-27B-Heretic-Antirep-V1 – AI Model by ToastyPigeon | AlphaNeural AI
You can deploy this model and start earning money today!
ToastyPigeon
/
Qwen3.5-27B-Heretic-Antirep-V1
like
0
safetensors
qwen3_5
qwen3.5
dpo
anti-repetition
uncensored
llmfan46/Qwen3.5-27B-ultra-uncensored-heretic-v1
finetune
other
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen3.5-27B-Heretic-Antirep-V1
Anti-repetition DPO fine-tune of
llmfan46/Qwen3.5-27B-heretic-v2
.
What is this?
This model applies a targeted DPO (Direct Preference Optimization) training pass to reduce degenerate repetition in long-form outputs while preserving the heretic model's uncensored behavior.
Training Details
Base model
:
llmfan46/Qwen3.5-27B-heretic-v2
Method
: QLoRA DPO (rank=32, alpha=16, RSLoRA)
Dataset
: 611 preference pairs (chosen=clean responses, rejected=repetitive responses)
Training
: 1 epoch, batch size 4, learning rate 5e-6, cosine schedule
Hardware
: 2× RTX 3090
LoRA targets
: All attention projections (GDN + standard) + MLP layers
Key Results
Eliminates catastrophic repetition loops on prompts that trigger degenerate output in the base model
Preserves the heretic model's uncensored/unrefused behavior — no new refusals introduced
Response quality and length remain comparable to the base model