Views
No views yet
19cdf7e)model.safetensors (~16 GB)
config.json
tokenizer.json
tokenizer_config.json
generation_config.json
chat_template.jinjacomfyui/qwen3-8b-heretic.safetensors # bf16, 16GB
comfyui/qwen3-8b-heretic_fp8_e4m3fn.safetensors # fp8, 8.8GB
comfyui/qwen3-8b-heretic_nvfp4.safetensors # nvfp4, 6.0GB| Quant | Size | Notes |
|---|---|---|
| F16 | 16GB | Lossless reference |
| Q8_0 | 8.2GB | Excellent quality |
| Q6_K | 6.3GB | Very good quality |
| Q5_K_M | 5.5GB | Good quality |
| Q5_K_S | 5.4GB | Slightly smaller Q5 |
| Q4_K_M | 5.0GB | Recommended balance |
| Q4_K_S | 4.8GB | Smaller Q4 variant |
| Q3_K_M | 3.9GB | For low VRAM only |
comfyui/qwen3-8b-heretic_fp8_e4m3fn.safetensors (8.8GB)comfyui/qwen3-8b-heretic_nvfp4.safetensors (6.0GB)comfyui/qwen3-8b-heretic.safetensors (16GB)ComfyUI/models/text_encoders/ClipLoader node and select the heretic file1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4model = AutoModelForCausalLM.from_pretrained(
5 "DreamFast/qwen3-8b-heretic",
6 device_map="auto",
7 torch_dtype=torch.bfloat16,
8 trust_remote_code=True
9)
10tokenizer = AutoTokenizer.from_pretrained("DreamFast/qwen3-8b-heretic")
11
12prompt = "Describe a dramatic sunset over a cyberpunk city"
13inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
14outputs = model.generate(**inputs, max_new_tokens=200)
15print(tokenizer.decode(outputs[0], skip_special_tokens=True))llama-server -m qwen3-8b-heretic-Q4_K_M.gguf? Which trial do you want to use?
[Trial 2732] Refusals: 10/100, KL divergence: 0.1001
> [Trial 2681] Refusals: 13/100, KL divergence: 0.0838 <-- selected
[Trial 2337] Refusals: 18/100, KL divergence: 0.0643
[Trial 2419] Refusals: 19/100, KL divergence: 0.0600
[Trial 2195] Refusals: 21/100, KL divergence: 0.0534
...