Views
No views yet
| Base model | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 |
| Format | W4A16 |
| Total params | 92B |
| Active / token | 12B |
| Experts / layer | 384 |
| Layers | — |
| Hidden size | 4096 |
| Context | 262,144 |
| On-disk size | 56 GB |
| Variant | Format | Link |
|---|---|---|
Nemotron-3-Super-64B | BF16 | link |
Nemotron-3-Super-64B-W4A16 | W4A16 | link |
Nemotron-3-Super-92B | BF16 | link |
Nemotron-3-Super-92B-W4A16 (this) | W4A16 | link |
/mnt/llm_models/nemotron-super-compressions/nemotron_super_merged_long50_short15120_v2/reap_25pctREAP 25% pruned checkpointintel/auto-round 0.10.2W4A16auto_roundNeelNanda/pile-10kauto12850102421True/home/ser/nemotron-super/autoround_w4a16/reap_25pct1@misc{lasby2025reap,
2 title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
3 author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
4 year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
5}