Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-coder-480b-folded-nvfp4-inputscale – AI Model by Soroosh-Sa | AlphaNeural AI
You can deploy this model and start earning money today!
Soroosh-Sa
/
qwen3-coder-480b-folded-nvfp4-inputscale
like
0
safetensors
qwen3_moe
8-bit
modelopt
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen3-Coder-480B-A35B Folded NVFP4 Checkpoint
This checkpoint was created from folded BF16 weights using a streaming tensor-wise NVFP4 converter.
Pipeline:
RMSNorm gamma folded into q/k/v and MLP gate/up-related weights.
Tensor-wise NVFP4 quantization using the official NVIDIA NVFP4 checkpoint as a layout template.
Fused QKV/gate-up weight_scale_2 groups normalized for TensorRT-LLM loading.
input_scale adjusted using original RMSNorm gamma statistics.
Validated:
Safetensors layout validation passed.
TensorRT-LLM PyTorch backend loads.
Smoke generation works on 8x GPUs.
Recommended runtime:
TensorRT-LLM 1.3.0rc16
transformers 5.5.4
TP_SIZE=8