Views
No views yet
compressed-tensors NVFP4A16 (E2M1 4-bit
weights, FP8-E4M3 scales @ group 16), for NVIDIA V100 (SM70) under 1Cat-vLLM.
Vision tower and the grafted MTP head (bf16) are preserved.apply_from=4). Fully cyber-open, general capability intact.| cyber-open ↑ | confab ↓ | factual ↑ | gsm8k ↑ | degen ↓ | |
|---|---|---|---|---|---|
| Cyber v2 (this line) | 100/100 | 0.867 | 1.00 | 0.80 | 0.00 |
| previous Cyber build | 93/100 | 1.00 | 0.933 | 0.825 | 0.00 |
[Linear], GDN in_proj_qkv/in_proj_z kept fp16, 768 CoT + 256 wiki calibration, actorder=weight, MSE observer. MTP head carried in bf16.--kv-cache-dtype fp8_e5m2, MTP speculative decoding. Serves on V100/SM70 via 1Cat prepare_nvfp4_linear (min capability 70).