Views
No views yet
nvfp4_mlp_only, MSE calibration on reasoning data). For high-throughput low-precision inference on NVIDIA Blackwell GPUs.| Item | Value |
|---|---|
| Format | NVFP4 (E2M1 + 2-level scale, group size 16), modelopt |
| Size | ~23.7 GB (3 shards) |
| Quantized | routed experts + shared-expert MLP → NVFP4 |
| Kept BF16 | attention, GatedDeltaNet linear-attn, MoE router, shared_expert_gate, lm_head, embeddings |
| KV-cache | not quantized (BF16) |
transformers (use a serving runtime).python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-codevllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code<think>…</think>). Official Qwen3.5 settings:temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0nvfp4_mlp_only、推理資料 MSE 校準)量化的 abliterated(去審查)Qwen3.5 35B MoE 推理模型,供 NVIDIA Blackwell GPU 高吞吐低精度推理。| 項目 | 數值 |
|---|---|
| 格式 | NVFP4(E2M1 + 兩級縮放,group size 16),modelopt |
| 大小 | ~23.7 GB(3 shards) |
| 量化 | routed experts + shared-expert MLP → NVFP4 |
| 保 BF16 | attention、GatedDeltaNet linear-attn、MoE router、shared_expert_gate、lm_head、embedding |
| KV-cache | 不量化(BF16) |
transformers 載入(請用推理 runtime)。python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-codevllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code<think>…</think>)。Qwen3.5 官方設定:temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0