Views
No views yet

| Field | Value |
|---|---|
| Source | poolside/Laguna-S-2.1 @ e80da38 |
| Architecture | laguna — 48 layers (12 global + 36 SWA w512), 118B-A8B, 256 experts top-10 + shared, 1M ctx |
| On-disk size | 96.5 GB (21 shards) |
| Routed experts | 6-bit gate/up/down affine, group 64, AWQ folded |
| Attention q/k/v/o + g_proj | 8-bit affine |
| Shared expert / dense FFN | 8-bit affine |
| Embeddings / lm_head | 6-bit / 8-bit affine |
| Router, e_score bias, norms | fp16 passthrough |
| Modality | text-only (verified from tensor index) |
| Metric | Value |
|---|---|
| Decode | 30.7 tok/s |
| Long-context cache parity | teacher-forced top-1 agreement 1.000 / 1.000 (pre/post the 512 sliding window, 2,913-token pass) |
enable_thinking toggles reasoning (thinking is ON by default in this revision's template; pass enable_thinking=False to disable)tokenizer_config.json (upstream ships only an {% include %} stub that most runtimes cannot resolve — inlining is what makes the reasoning toggle actually work)eos_token_id = [2, 24] — id 24 is end-of-turn and must be in the stop set〈|EOS|〉 (bos 2): do not prepend another<tool_call>name<arg_key>k</arg_key><arg_value>v</arg_value></tool_call>{bits, group_size, mode} overrides in config.json[quantization].