Views
No views yet
| File | Size | Description |
|---|---|---|
llm.mnn | 8.5MB | Model graph |
llm.mnn.weight | 14GB | Quantized weights |
embeddings_bf16.bin | 2.4GB | BF16 embedding table (required) |
llm_config.json | 4.4KB | Model config with jinja chat template |
tokenizer.txt | 2.9MB | Tokenizer |
Qwen3.5 hybrid LinearAttention models are usually better on CPU than OpenCL in current TokForge routing.