Views
No views yet
tokenizer.json so the volume is self-contained.mmproj-* files (~160 GB
total) and no tokenizer.json. This volume carries ONLY the one served quant
plus the tokenizer (same rationale as qwen2.5-0.5b-gguf / qwen3.5-122b-gguf):
a multi-gguf volume makes the host's preload pick ambiguous, and wrapping the
whole repo wastes ~154 GB of attested volume for files that never serve. The
mmproj files are omitted deliberately — llm-chat's ggml path is text-only and
preloads a single LLM gguf.| File | Upstream | Revision | sha256 |
|---|---|---|---|
Qwen3.5-9B-UD-Q4_K_XL.gguf | unsloth/Qwen3.5-9B-GGUF | 3885219b6810b007914f3a7950a8d1b469d598a5 | 6f5d30666c2d8ae16a306e616d95341dcf3cc46810df84d7e6f5a7d1e4c1b293 |
tokenizer.json | Qwen/Qwen3.5-9B | c202236235762e1c871ad0ccb60c8ee5ba337b9a | 5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42 |
qwen3.5-122b-gguf (the Qwen3.5
family shares it — same sha256, same vocab 248320, <|im_end|> 248046 /
<|endoftext|> 248044).LICENSE is the Apache-2.0 text; both upstreams are Apache-2.0./models/<name>; the host preloads the GGUF as the wasi-nn ggml graph. As the
volume's single gguf it needs no MODEL_VOLUMES third field and no model_file
override, and with the tokenizer bundled, llm-chat's default tokenizer.json
lookup works with no cross-volume configuration.