[384 of 512 routed experts kept per layer - 25% of experts pruned]
from
inclusionAI/Ling-3.0-flash
(124B total / 5.1B active).
Method: one-shot REAP (
Router-weighted Expert Activation Pruning) -
experts scored by router-gate-value × output-L2-norm over calibration data, lowest-scoring deleted.
No fine-tuning, no recovery training.
BF16 safetensors. Loads with trust_remote_code=True (custom bailing_hybrid / BailingMoeV3 code).
Research artifact - quantized builds live in the sibling -GGUF repo.