NVFP4A16 (W4A16: FP4 weights, BF16 activations; compressed-tensors) quantization of
upstage/Solar-Open2-250B
(250B-A15B MoE, 320+1 experts, KDA linear attention, 1M context).
Quantized by
Lna-Lab (
@Tono_Ken3) on 12x RTX PRO 2000 Blackwell / TR PRO 9985WX,
llm-compressor 0.12 + upstage transformers fork (v5.14.1-solar-open2).
Our focus here is a fully documented conversion: every trap we hit is written down.
Expert weights use per-expert (unfused) names. Serve with the upstage vLLM fork;
a fused-name loader may require a weight-name mapping shim (notes to follow).
MEASURED: TP=4 x PP=3 on 12x RTX PRO 2000 (16GB, 60W cap): single-stream 64.9 t/s, C8 aggregate 217.8 t/s. Probes: exact answers on number-theory and AIME-style tasks.
Upstage Solar License (Apache-2.0-based; commercial use permitted). Per Section 4(e),
this derivative keeps the "Solar" brand in its name.
Our first bake used full NVFP4 W4A4 (48x2048 CPU calibration, basic pipeline).
That build loaded and ran at full speed but emitted only empty/special tokens;
requantizing weights-only (NVFP4A16) with the identical calibration was fully
stable. We report this as an observation under OUR calibration recipe, not a
property of the architecture — W4A4 may well be achievable with different
calibration choices.