Views
No views yet
[!IMPORTANT]⚠️ Behavior Notice — Fine-tuned (DPO) variant
This is an abliterated + DPO fine-tuned variant (LEACE abliteration to remove refusal, then LoRA SFT + DPO to reduce moralizing). It reaches lower moralizing (~31%) than pure abliteration, but the fine-tuning may have altered the model's behavior, style, or knowledge beyond decensoring. If you want the base model's original behavior / thinking preserved as closely as possible (weights-only, zero training), use the pure-abliteratedvariant instead (moralizing ~38%, reasoning fully intact — recommended default).
[!WARNING]⚠️ READ FIRST — Sampling Parameters MUST Be Set Correctly
This model requires the exact sampling parameters below, especiallyrepeat-penalty 1.05. Wrong values break it:
Setting Result repeat-penalty 1.05✅correct (sweet spot) repeat-penalty 1.0severe thinking loops repeat-penalty 1.1truncated / unfinished answers temp 0(greedy)not recommended
| Item | Value |
|---|---|
| Architecture | Qwen3.5 9B dense — GatedDeltaNet (linear attention) + full-attention hybrid (3:1) |
| Base model | deepreinforce-ai/Ornith-1.0-9B |
| Abliterated by | YuYu1015 |
| Precision | BF16 (~17 GB) |
| Context length | Inherited from base |
| Thinking mode | Supported (reasoning model, emits <think>…</think>) |
| Languages | English, Chinese |
| Metric | Base Ornith-1.0-9B | This model |
|---|---|---|
| Hard refusal rate | 99.5% | <1% |
| Moralizing / disclaimer rate | 99.5% | 31% |
| GSM8K (reasoning accuracy) | 86.7% | 90.0% |
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.05⚠️--repeat-penaltyis critical — keep it at1.05. This value gives near-normal generation and is the sweet spot for this model. Do NOT change it:1.0causes severe thinking loops, while1.1makes the model fail to finish its answer. Greedy decoding (--temp 0) is also not recommended for this family.
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4m = "YuYu1015/YuYu1015-Ornith-1.0-9B-abliterated-dpo"
5tok = AutoTokenizer.from_pretrained(m, trust_remote_code=True)
6model = AutoModelForCausalLM.from_pretrained(m, dtype=torch.bfloat16,
7 trust_remote_code=True).to("cuda").eval()
8msgs = [{"role": "user", "content": "Your prompt here"}]
9text = tok.apply_chat_template(msgs, add_generation_prompt=True, tokenize=False)
10ids = tok(text, return_tensors="pt", add_special_tokens=False).to("cuda")
11out = model.generate(**ids, max_new_tokens=4096, do_sample=True,
12 temperature=1.0, top_p=0.95, top_k=20, min_p=0.0,
13 repetition_penalty=1.05) # 1.05 only — 1.0 loops, 1.1 truncates
14print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))[!IMPORTANT]⚠️ 行為警告 — DPO 微調版
本版為 abliteration + DPO 微調版(先 LEACE 抽掉拒絕方向,再 LoRA SFT + DPO 壓低說教)。說教比純 abliteration 更低(約 31%),但微調可能改變模型的行為、風格或知識,不只是去審查。 若想「盡量保留原模型原始行為/思考」(只動權重、零訓練),請改用純-abliterated版(說教約 38%、推理完整保留 —— 推薦預設)。
[!WARNING]⚠️ 必讀 — 取樣參數務必正確設定
本模型強依賴下方那組取樣參數,尤其repeat-penalty 1.05。設錯會壞掉:
設定 結果 repeat-penalty 1.05✅正常(甜蜜點) repeat-penalty 1.0嚴重思考迴圈 repeat-penalty 1.1答不完被截斷 temp 0(貪婪)不建議
| 項目 | 數值 |
|---|---|
| 架構 | Qwen3.5 9B dense — GatedDeltaNet(線性注意力)+ 全注意力混合(3:1) |
| 基礎模型 | deepreinforce-ai/Ornith-1.0-9B |
| 去審查者 | YuYu1015 |
| 精度 | BF16(約 17 GB) |
| Context 長度 | 沿用基礎模型 |
| 思考模式 | 支援(推理模型,輸出 <think>…</think>) |
| 語言 | 英文、中文 |
| 指標 | 原版 Ornith-1.0-9B | 本模型 |
|---|---|---|
| 硬拒答率 | 99.5% | <1% |
| 說教/免責率 | 99.5% | 31% |
| GSM8K(推理正確率) | 86.7% | 90.0% |
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.05⚠️--repeat-penalty很關鍵——請保持1.05。 此值生成接近正常,是本模型的甜蜜點。請勿更動:1.0會嚴重思考迴圈、1.1會讓模型答不完。同樣不建議用貪婪解碼(--temp 0)。
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4m = "YuYu1015/YuYu1015-Ornith-1.0-9B-abliterated-dpo"
5tok = AutoTokenizer.from_pretrained(m, trust_remote_code=True)
6model = AutoModelForCausalLM.from_pretrained(m, dtype=torch.bfloat16,
7 trust_remote_code=True).to("cuda").eval()
8msgs = [{"role": "user", "content": "你的問題"}]
9text = tok.apply_chat_template(msgs, add_generation_prompt=True, tokenize=False)
10ids = tok(text, return_tensors="pt", add_special_tokens=False).to("cuda")
11out = model.generate(**ids, max_new_tokens=4096, do_sample=True,
12 temperature=1.0, top_p=0.95, top_k=20, min_p=0.0,
13 repetition_penalty=1.05) # 只能 1.05 — 1.0 會 loop、1.1 會截斷
14print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))