Views
No views yet
| Check | Original Qwen3.6-27B | This MLX Quant |
|---|---|---|
| Official 25-prompt refusal check | 20/25 refusals | 0/25 refusals |
| 100-prompt refusal check | 92/100 refusals | 3/100 refusals |
Youssofal/Qwen3.6-27B-Abliterated-Heretic-Uncensored-BF16 using dynamic layer-aware 3/4/5-bit affine MLX quantization. It is a dynamic, layer-aware quantization pass rather than a flat conversion. The retained multimodal payload includes the source vision/video tensors and processor files; direct image/video use depends on MLX runtime support for Qwen3.6 multimodal inputs.model-*.safetensors: quantized MLX text weightsmodel-vision-00001-of-00001.safetensors: retained BF16 vision/video tensor shardmodel.safetensors.index.json: tensor-to-shard mapping, including retained vision tensorsconfig.json, generation_config.json, configuration.json: model configtokenizer.json, tokenizer_config.json, chat_template.jinja: tokenizer + chat templatepreprocessor_config.json, video_preprocessor_config.json: image + video processor configsmlx_variant_metadata.json: build metadata1from mlx_lm import load, generate
2from mlx_lm.sample_utils import make_sampler
3
4repo = "Youssofal/Qwen3.6-27B-Abliterated-Heretic-Uncensored-MLX-3bit"
5model, tokenizer = load(repo)
6
7messages = [{"role": "user", "content": "Write a short Python function that reverses a string."}]
8prompt = tokenizer.apply_chat_template(
9 messages,
10 tokenize=False,
11 add_generation_prompt=True,
12 enable_thinking=False,
13)
14
15response = generate(
16 model,
17 tokenizer,
18 prompt=prompt,
19 max_tokens=256,
20 sampler=make_sampler(temp=0.0),
21)
22print(response)| Spec | Value |
|---|---|
| Total Parameters | 27.8B (dense source) |
| Layers | 64 |
| Attention | Hybrid (3 linear-attention + 1 full-attention per 4-layer group) |
| Hidden Size | 5120 |
| Family | qwen3_5 |
| Modality | Vision-language source; MLX text path locally validated |
| MLX Runtime Validation | Text generation |
| Base Model | Qwen/Qwen3.6-27B |