Views
No views yet
Important: This model uses the JANG quantization format — the GGUF equivalent for MLX on Apple Silicon. Currently only supported by MLX Studio and thejang-toolsPython package.

| Architecture | Qwen 3.5 VL Dense — 27B params, hybrid SSM/FA, 64 layers |
| Quantization | JANG_4S (6/4-bit mixed) — 16 GB |
| Abliteration | CRACK — novel weight surgery |
| HarmBench | 75.0% (240/320) |
| MMLU | 83.1% (base: 83.1%, 0% drop) |
| Speed | 27 tok/s (M4 Max) |
| Vision | Yes — via MLX Studio / vMLX |
| Thinking | ON/OFF supported |
| Fits on | 32 GB+ Macs |
| Model | MMLU | Size | Speed | Notes |
|---|---|---|---|---|
| JANG_4S + CRACK | 83.1% | 16 GB | 27 tok/s | This model |
| JANG_4S (base) | 84.5% | 16 GB | 35 tok/s | Unmodified JANG |
| MLX 4-bit | 84.5% | 14 GB | 20 tok/s | Uniform quant |
| MLX 8-bit | ~86% | 29 GB | ~15 tok/s | 2x larger |
enable_thinking=false, temperature=1.0| Category | Score | |
|---|---|---|
| Misinformation / Disinfo | 47/54 | 87% |
| Copyright | 68/80 | 85% |
| Chemical / Biological | 35/42 | 83% |
| Illegal | 38/53 | 72% |
| Harmful | 12/18 | 67% |
| Cybercrime / Intrusion | 31/52 | 60% |
| Harassment / Bullying | 9/21 | 43% |
Note: Dense models have stronger distributed safety training than MoE models, making them harder to fully abliterate while preserving knowledge. This model prioritizes zero MMLU degradation over maximum compliance.
| Subject | CRACK | Base | Delta |
|---|---|---|---|
| College Physics | 5/5 | 5/5 | 0 |
| Professional Medicine | 5/5 | 5/5 | 0 |
| Conceptual Physics | 5/5 | 5/5 | 0 |
| Electrical Engineering | 5/5 | 5/5 | 0 |
| Machine Learning | 5/5 | 5/5 | 0 |
| HS Biology | 5/5 | 5/5 | 0 |
| Abstract Algebra | 4/5 | 4/5 | 0 |
| College CS | 4/5 | 4/5 | 0 |
| HS Geography | 4/5 | 4/5 | 0 |
| World Religions | 5/5 | 5/5 | 0 |
| HS Mathematics | 3/5 | 3/5 | 0 |
| Formal Logic | 3/5 | 3/5 | 0 |
| College Math | 1/5 | 1/5 | 0 |
| Total | 54/65 (83.1%) | 54/65 (83.1%) | 0% |
pip install "jang[mlx]"1from jang_tools.loader import load_jang_model
2from mlx_lm import generate
3
4model, tokenizer = load_jang_model("dealignai/Qwen3.5-VL-27B-JANG_4S-CRACK")
5
6messages = [{"role": "user", "content": "Your prompt here"}]
7prompt = tokenizer.apply_chat_template(
8 messages, add_generation_prompt=True, tokenize=False)
9
10response = generate(model, tokenizer, prompt=prompt, max_tokens=2000)
11print(response)1prompt = tokenizer.apply_chat_template(
2 messages, add_generation_prompt=True,
3 enable_thinking=False, tokenize=False)Tip: Usetemperature=1.0for chat (greedy can cause repetition). Usetemperature=0.0for structured tasks like MMLU.
| 항목 | 내용 |
|---|---|
| 크기 | 16 GB |
| HarmBench | 75.0% (240/320) |
| MMLU | 83.1% (기본 대비 0% 하락) |
| 속도 | 27 tok/s (M4 Max) |
| 비전 | 지원 (MLX Studio / vMLX) |
| 최소 요구사양 | 32 GB 메모리 Mac |
pip install "jang[mlx]"