Views
No views yet

[!NOTE] These NVFP4s are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
| Property | Value |
|---|---|
| Base model | poolside/Laguna-S-2.1 |
| Parameters | 117.6B |
| Layers | 48 |
| Experts | 256 routed (top-10) |
| Sliding window | 512 tokens |
| Context length | 1,048,576 tokens (1M) |
| Vocabulary | 100,352 |
| Modalities | Text |
| Architecture | Mixture-of-Experts, 256 experts (top-10), hybrid sliding-window (512) and global attention, 48 attention heads over 8 KV heads, LagunaForCausalLM |
| This repo | NVFP4 weights |
| Benchmark | Score |
|---|---|
| Laguna S 2.1 | 70.2% |
| Tencent Hy3 | 71.7% |
| Inkling | 63.8% |
| Nemotron 3 Ultra | 56.4% |
| DeepSeek-V4-Pro Max | 64.0% |
| Kimi K3 | 88.3% |
| Qwen 3.7 Max | 74.5% |
| Muse Spark 1.1 | 80% |
| Claude Fable 5 | 88% |
poolside/Laguna-S-2.1, not our own measurements. Quantization preserves the large majority of this; Q4_K_M and up stay close to full precision.AtomicChat/Laguna-S-2.1-NVFP4 and hit Use this model.vllm serve AtomicChat/Laguna-S-2.1-NVFP4 --max-model-len 8192| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 1.0 |
| top_k | 20 |
| min_p | 0.0 |
poolside/Laguna-S-2.1.poolside/Laguna-S-2.1 (original weights).llm-compressor over a calibration corpus.