Views
No views yet

qwen3_5_moe). Quantized by Osaurus; runs on Apple Silicon via Osaurus / mlx_lm.vision_config with no vision-tower weights, so this bundle is correctly stamped modality: text (has_vision: false).| Family | qwen3_5_moe (hybrid) |
| Layers | 40 — 30 linear-attention (Gated DeltaNet / SSM) + 10 full-attention (1 every 4) |
| Experts | 256 routed (8 active) |
| Active params | ~3 B |
| Cache | hybrid (recurrent GDN state + KV for full-attn layers) |
conv1d / A_log / dt_bias / gated in_proj_{qkv,a,b,z}) carry a recurrent state across the full-attention layers — verified coherent with long-context recall in the Osaurus (vMLX-Swift) runtime.python -m mlx_lm generate --model OsaurusAI/Qwen-AgentWorld-35B-A3B-JANG_4M --prompt "Explain a hash map in two sentences."qwen tool parser).