Views
No views yet
Qwen/Qwen3.6-35B-A3B.
The checkpoint was produced with routed-expert pruning using REAP
(Router-weighted Expert Activation Pruning), which scores routed experts
with router weights and expert activation norms.| Setting | Value |
|---|---|
| Base model | Qwen/Qwen3.6-35B-A3B |
| Compression / pruning ratio | 0.50 |
| Pruning method | reap |
| Calibration samples | 1024 |
| Calibration sequence length | 2048 |
| Seed | 42 |
| Router weight renormalization | true |
| Routed experts per MoE layer | 256 -> 128 |
| Routed experts selected per token | 8 |
| Shared experts | Preserved |
| Precision | BF16 |
| Quantization | None |
theblackcat102/evol-codealpaca-v1: 171 samplesSalesforce/xlam-function-calling-60k: 171 samplesopen-r1/Mixture-of-Thoughts[code]: 171 samplesopen-r1/Mixture-of-Thoughts[math]: 171 samplesopen-r1/Mixture-of-Thoughts[science]: 170 samplesSWE-bench/SWE-smith-trajectories(tool): 170 samplesqwen3_5_moe architecture and includes tokenizer and
processor files.1@inproceedings{
2 lasby2026reap,
3 title={{REAP} the Experts: Why Pruning Prevails for One-Shot MoE compression},
4 author={Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
5 booktitle={The Fourteenth International Conference on Learning Representations},
6 year={2026},
7 url={https://openreview.net/forum?id=ukGxWd2aDG}
8}