Ling-3.0-flash is inclusionAI's next-generation native hybrid reasoning model:
124B total / 5.1B active (~12.4% of their previous 1T-class flagship). It uses a native hybrid-linear stack from pretraining —
5:1 Kimi Delta Attention (KDA) + gated MLA, 1/64 sparse MoE, 512 routed experts (top-8) + 1 shared expert, 2 dense layers, and a trained
MTP head. Official context schedule is 8K → 32K → 256K. Thinking is on by default in the official card.