Views
No views yet
Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP4, which is an NVFP4
compressed-tensors quantized version of
RangerX/Qwen3.6-35B-REAP-Pruned-ratio-0.5.Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP4 Hugging Face checkpoint with a patched
llama.cpp converter that handles compressed-tensors NVFP4 tensors for this
model.1uv run /home/sroecker/src/llama.cpp/convert_hf_to_gguf.py \
2 --verbose \
3 --outtype auto \
4 --outfile Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP4-GGUF/Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP4.gguf \
5 Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP439. The resulting GGUF contains 1293 tensors:280 NVFP4 tensors152 BF16 tensors861 F32 tensorsknoopx/Qwen3.6-35B-A3B-NVFP4-GGUF.
The converter-side changes used here follow the approach from the open
llama.cpp PR
#21095, which adds
conversion support for Hugging Face NVFP4 models quantized with
compressed-tensors. Native Blackwell NVFP4 CUDA runtime support is tracked
separately in #22196.compressed-tensors checkpoints with format nvfp4-pack-quantized as
NVFP4 inputsweight_packed, weight_global_scale,
and input_global_scale to the ModelOpt-style names expected by the
repacker1444df51289cfa8063b96f0e62b1125440111bc79a52003ea14b6eac7016fd5f as
qwen35| File | Size |
|---|---|
Qwen3.6-35B-REAP-Pruned-ratio-0.5-NVFP4.gguf | 13G |