Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Qwen3.6-35B-A3B-MTP-GGUF – AI Model by saidonnet | AlphaNeural AI
You can deploy this model and start earning money today!
saidonnet
/
Qwen3.6-35B-A3B-MTP-GGUF
like
0
gguf
endpoints_compatible
us
conversational
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen3.6-35B-A3B-MTP-GGUF (Quantized)
Overview
Quantized GGUF versions of Qwen3.6-35B-A3B-MTP (Apache MoE) for efficient local inference.
Source
Original
:
ggml-org/Qwen3.6-35B-A3B-MTP-GGUF
Base Model
: Qwen3.6-35B-A3B (Apache MoE, 35B parameters)
BF16 Source
: 67GB downloaded from ggml-org (official conversion, 1200+ downloads)
Available Quantizations
File
Size
Quantization
Quality
Qwen3.6-35B-A3B-MTP-Q4_0.gguf
19GB
Q4_0
Good
Qwen3.6-35B-A3B-MTP-Q4_K_S.gguf
19GB
Q4_K_S
Better
Quantization Details
Date
: 2026-07-25
Tool
: llama.cpp (commit b1-fb92d8f)
VM
: Lightning.ai (16 cores, 60GB RAM)
Commands
:
Usage
llama-server
llama-cli
Features
MTP Support
: MoE Targeted Prediction for faster inference
A3B Architecture
: Apache 3-Branch MoE design
35B Parameters
: Full model capacity in quantized form
GGUF Format
: Compatible with llama.cpp, llama-server, llama-cli
License
Same as original:
Apache 2.0
Credits
Original model: Qwen team
GGUF conversion: ggml-org
Quantization: saidonnet (2026-07-25)