Views
No views yet
| Metric | FP16 | v1 (Q5) | v3 |
|---|---|---|---|
| Download | 69.3 GB | 35.2 GB | 15.6 GB |
| PPL | 5.19 | 5.39 | 5.36 |
| Compression | 1.0x | 2.0x | 4.44x |
| tok/s | 30.1 | 30.2 | 30.1 |
| GPU dequant | --- | ~400s | 8.8s |
| Tensor Type | Bits |
|---|---|
| Expert gate_up_proj | Q3 |
| Expert down_proj | Q4 |
| Attention Q/K/V | Q5 |
| Attention O | Q6 |
| MoE router, norms | FP16 |
1from huggingface_hub import snapshot_download
2import sys
3local = snapshot_download("caiovicentino1/Qwen3.5-35B-A3B-EOQ-v3")
4sys.path.insert(0, local)
5from eoq_loader import load_eoq_model
6model, tokenizer = load_eoq_model("caiovicentino1/Qwen3.5-35B-A3B-EOQ-v3")