
compressed-tensors / pack-quantized formatlm_head remain BF16model.safetensors file| Qwen 3.8 27B | Focus-Red-Int4 | |
|---|---|---|
| author | Alibaba Qwen | Jaid |
| repository | Qwen/Qwen3.8-27B | Jaidchen/Focus-Red-Int4 |
| architecture | qwen3_5 | qwen3_5_text |
| Transformers handler |
Qwen3_5ForConditionalGeneration
|
Qwen3_5ForCausalLM
|
| tensor entries | 1199 | 2051 |
| tensor type | bf16 | W4A16 G32 asymmetric + selected BF16 |
| parameters | 27 781 427 952 | 26 895 998 464 |
| active | 100% | 100% |
| vocabulary size | 248 320 | 248 320 |
| context size | 262 144 | 262 144 |
| MTP | integrated | detached → Focus-Red-MTP |
| sampling strategy | random sampling | greedy/deterministic |
| sampling parameters |
do_sample: true
temperature: 1.0 top_k: 20 top_p: 0.95 |
do_sample: false
temperature: 0 top_k: 1 top_p: 1 |
| input modality | text, image, video | text |
| model size | 55 562 855 904 | 19 202 352 336 bytes on disk |
| splits | 18 | none |
| Jinja template | Qwen original | focus-chat-template dist build |
compressed-tensorspack-quantizedcompressedconsult tool to your harness that calls a vision-enabled subagent model like Gemini Flash or GPT.Qwen3_5ForCausalLM class which may not be included in your inference engine. In this case you would need to ask your agent or integrate it yourself.
dist/chat_template.jinja from jaidlab/focus-chat-template. The template is reproducibly built from Qwen/Qwen3.8-27B's pinned upstream template at commit 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 plus the repository's ordered patch stack.5c381ca45e9538c7a2406331b554ee7d62cf3d0b8c115f17687d4fdd5590a239