[!TIP]
This model is reproducible!
See the
README in the
reproduce directory for more information.
Abliteration parameters
| Parameter | Value |
|---|
| direction_index | 20.01 |
| attn.o_proj.max_weight | 1.50 |
| attn.o_proj.max_weight_position | 20.75 |
| attn.o_proj.min_weight | 1.50 |
| attn.o_proj.min_weight_distance | 17.12 |
| mlp.down_proj.max_weight | 1.24 |
| mlp.down_proj.max_weight_position | 19.56 |
| mlp.down_proj.min_weight | 1.16 |
| mlp.down_proj.min_weight_distance | 11.01 |
Performance
| Metric | This model | Original model (empero-ai/Qwen3.8-4B) |
|---|
| KL divergence | 0.0167 | 0 (by definition) |
| Refusals | 6/100 | 99/100 |
Qwen3.8-4B
[!Note]
This repository contains model weights and configuration files in the Hugging Face Transformers format.
These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, and other standard runtimes with Qwen3.5 architecture support.
Qwen3.8-4B is a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture. The student was trained on ~45,000 curated teacher traces from our internal Qwen3.8 distillation datasets — dense chain-of-thought spanning mathematics, general reasoning, and instruction following, quality-filtered before training.
The objective: bring the reasoning behavior of a frontier-scale teacher into a 4B that runs comfortably on consumer hardware.
Highlights
- Distilled chain-of-thought — every answer opens with a
<think> block learned directly from Qwen3.8 2.4T A95B traces rather than synthetic self-generated reasoning.
- 4B weight class — bf16 fits in ~8 GB; quantized builds run on laptops and consumer GPUs.
- Native function calling per Qwen3.5's specification — no wrapper or tool-specific fine-tune required.
- 262,144-token native context, inherited from the Qwen3.5 base.
- Full fine-tune — every parameter updated; not an adapter.
Model Overview
- Type: Causal Language Model (text path of a vision-language base)
- Base: Qwen/Qwen3.5-4B
- Number of Parameters: 4B
- Training: SFT (off-policy distillation) on ~45,000 teacher traces
- Teacher: Qwen3.8 2.4T A95B (internal distillation datasets)
- Context Length: 262,144 natively
Benchmark Results
Measured with
lm-evaluation-harness, HF backend, identical settings for base and student. Both models are reasoning models and are evaluated with the CoT protocols (
gsm8k_cot,
mmlu_flan_cot_zeroshot); MMLU covers all 57 subjects (~1,700 questions). Flexible-extract is the primary metric; strict-match requires exact answer formatting.
| Task | Metric | Qwen3.5-4B (base) | Qwen3.8-4B | Δ |
|---|
| gsm8k_cot | exact_match (flexible) | 0.850 | 0.785 | −0.065 |
| gsm8k_cot | exact_match (strict) | 0.850 | 0.785 | −0.065 |
| mmlu (CoT, 57 subjects) | acc (flexible-extract) | 0.354 | 0.553 | +0.199 |
| mmlu (CoT, 57 subjects) | acc (strict-match) | 0.071 | 0.233 | +0.162 |
Sampling for generation: temperature=0.6, top_p=0.95, top_k=20 (Qwen3.5 recommended settings).
Quickstart
1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4model_id = "empero-ai/Qwen3.8-4B"
5tok = AutoTokenizer.from_pretrained(model_id)
6model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
7
8messages = [{"role": "user", "content": "A snail is at the bottom of a 10-meter well. Each day it climbs 3 meters, each night it slips back 2. How many days until it escapes?"}]
9inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
10
11out = model.generate(inputs, max_new_tokens=16384,
12 temperature=0.6, top_p=0.95, top_k=20, do_sample=True)
13print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
A recent
transformers release with Qwen3.5 support is required, along with the Gated DeltaNet kernels (
flash-linear-attention and a CUDA-matched
causal_conv1d build) — without them the linear-attention layers fall back to slow, memory-hungry PyTorch ops.
Best Practices
- Sampling:
temperature=0.6, top_p=0.95, top_k=20. Greedy decoding on long generations is a known repetition-loop failure mode for reasoning models in this class.
- Output length: allow generous
max_new_tokens (16,384 recommended); every answer opens with a <think> block. Parse and strip the <think>...</think> span for end users.
- Scope: the trace mix emphasizes mathematics, reasoning, and instruction following; for the strongest code performance in the family, use Qwen3.8-9B. The fine-tune is text-only; vision behavior is inherited from the base and was not evaluated here.
Stay in the loop
Sign up for the Empero newsletter at
empero.org for releases, evals, and research notes.
Support / Donate
If this model helped you, consider supporting the project:
- BTC:
bc1qx6zepu6sfkvshgdmc4ewu6pk6rpadvpgffpp7v
- LTC:
ltc1qv2mefzps2vtjcpwfx8xxdrpplrcvltswm68r7x
Provenance & licensing
Weights are released under Apache-2.0, inherited from the Qwen3.5-4B base. Shared for research and experimentation, as-is.
Acknowledgements
- Developed and released by Empero
- Base model: Qwen3.5-4B (Alibaba Qwen team)
- Training: TRL + Transformers
- Linear-attention kernels: flash-linear-attention, causal_conv1d
- Evaluation: lm-evaluation-harness (EleutherAI)