bringing warmth, emotional intelligence, general chat improvement to Qwen 3.5 series
countering some negative tendencies of Heretic models (overwillingness to agree, be sycophantic, etc)
This is still intended as a general use model (agentic, coding, general chat). Tuning was lightly & with precision. More general benchmarks to follow.
What this model does
This model is trained to be a better conversational partner in emotionally complex situations, while maintaining base model capabilities. It:
Validates without sycophancy — empathizes with frustration without rubber-stamping bad behavior
Sets boundaries warmly — names uncomfortable truths without lecturing
Sounds human — conversational tone, not therapist-speak. better tone vs vanilla Qwen 3.5, e.g. "It sounds like"
Key specs
Base
Qwen/Qwen3.5-35B-A3B
Parent
llmfan46/Qwen3.5-35B-A3B-heretic-v2 (decensored via MPOA+SOMA)
Fine-tune
DPO with LoRA (r=32, alpha=64)
Training data
DPO preference pairs with diverse, simulated (real-situation-based) generated dialogue
EQ-Bench 3 results
Ranked #8 on raw score*EQ-Bench 3 with only 3B active parameters — competitive with frontier models at a fraction of the compute.
#
Model
Raw Score
1
horizon-alpha
202.3
2
Kimi-K2-Instruct
202.0
3
gemini-2.5-pro-preview-06-05
200.5
4
o3
199.0
5
gpt-5
195.6
8
EQ-v5 (this model, 3B active)
193.6
10
claude-opus-4
192.6
*Table lists all models available in EQ-Bench 3 repo (so known judge, settings etc so we can be as apples to apples on raw score). Still raw score is not ideal. ELO submission pending. Better than no stats!