Views
No views yet
black-forest-labs/FLUX.1-schnellnunchaku-ai/nunchaku-flux.1-schnellsvdq-fp4_r32-flux.1-schnell.safetensorsquant_method: nunchaku_lite, NVFP4 SVDQ with group size 16, runtime rank 64, 418 SVDQ targets, and 76 AWQ W4A16 targets. The CLIP encoder is copied from the base model and T5 text_encoder_2 is BitsAndBytes 4-bit NF4. Fused QKV modules are split in logical tensor layout; single-block proj_out is merged from attention and MLP projections; low-rank tensors are logically padded to rank 64. NVFP4 outer scales are reconciled without overflowing FP8 group scales.| Checkpoint | Latency | Max VRAM |
|---|---|---|
| Converted Diffusers Nunchaku Lite NVFP4 r32 + BNB4 T5 | 1.34 s (stdev 0.03 s) | 16.73 GiB |
| Original Nunchaku NVFP4 r32 + BF16 T5 | 0.81 s (stdev 0.00 s) | 20.99 GiB |
nvidia-smi, including allocations outside PyTorch's caching allocator. The native row uses the base model's BF16 T5 encoder.
kernels package and a Blackwell NVIDIA GPU for NVFP4 kernels.1import torch
2from diffusers import FluxPipeline
3
4pipe = FluxPipeline.from_pretrained(
5 "lite-infer/flux.1-schnell-nunchaku-lite-nvfp4_r32-bnb4-text-encoder",
6 torch_dtype=torch.bfloat16,
7).to("cuda")
8
9image = pipe(
10 prompt='A cinematic photograph of a red fox standing in a misty forest at sunrise, detailed fur, volumetric light',
11 generator=torch.Generator("cuda").manual_seed(0),
12 width=1024,
13 height=1024,
14 num_inference_steps=4,
15 guidance_scale=0.0,
16).images[0]
17image.save("output.png")