Views
No views yet
cerebras/GLM-4.7-Flash-REAP-23B-A3B. This architecture implements Cerebras's Router-weighted Expert Activation Pruning, which mathematically identifies and removes the 16 least important experts from the original 64-expert frame.| Model Variant | Active Experts | Quant | Logic (GSM8K PPL) | Result |
|---|---|---|---|---|
| Viveka-23B (This Model) | 48 | Q6 | 3.23 | 👑 THE CHAMPION |
| Original GLM-4.7 (Full) | 64 | Q8 | 3.28 | Baseline Target |
| REAP Base (Pruned Only) | 48 | Q8 | 3.30 | Pruning Cost: 0.6% |
| Viveka-23B (Low-bit) | 48 | Q4 | 9.50 | Total Logic Failure |
mlx-lm and run:1from mlx_lm import load, generate
2
3model, tokenizer = load("your-hf-username/Viveka-GLM-4.7-23B-REAP-Smarty-MLX")
4response = generate(model, tokenizer, prompt="Solve: If a cube has a side of 4, what is the surface area?", verbose=True)