Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
cerebras_GLM-4.5-Air-REAP-82B-A12B-GGUF – AI Model by TobDeBer | AlphaNeural AI
You can deploy this model and start earning money today!
TobDeBer
/
cerebras_GLM-4.5-Air-REAP-82B-A12B-GGUF
like
0
transformers
glm
moe
pruning
compression
text-generation
en
zai-org/GLM-4.5-Air
finetune
mit
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
📋 Model Overview
GLM-4.5-Air-REAP-82B-A12B
has the following specifications:
Base Model
: GLM-4.5-Air
Compression Method
: REAP (Router-weighted Expert Activation Pruning)
Compression Ratio
: 25% expert pruning
Type
: Sparse Mixture-of-Experts (SMoE) Causal Language Model
Number of Parameters
: 82B total, 12B activated per token
Number of Layers
: 46
Number of Attention Heads (GQA)
: 96 for Q and 8 for KV
Number of Experts
: 96 (uniformly pruned from 128)
Number of Activated Experts
: 8 per token
Context Length
: 131,072 tokens
License
: MIT