This is the next-level face-swap specialized evolution of the Dark Beast lineage, built on the lightning-fast FLUX.2 Klein 9B accelerated model from Black Forest Labs.
Engineered with targeted optimizations for face swapping practices, it integrates BFS (Best Face Swap) technology to completely eliminate the rigid, unnatural look that plagued earlier face replacements — delivering seamless, lifelike integrations with preserved identity, expression, and lighting.
It also fully fixes the portrait reference issue from the previous DB BlitZ versions, ensuring right reference adherence every time.
Special thanks to the scheme provider: https://github.com/alisson-anjos for the powerful BFS foundation that powers this breakthrough.🟦
image
Important notes:
This version is exclusively designed around the Klein 9B accelerated edition — no base model exists.
Usage is identical to Black Forest Labs' official FLUX.2 Klein 9B accelerated release: ultra-low steps (e.g., 4-5), CFG=1 fixed, blazing inference speed on consumer hardware.
In one sentence: Dark Beast's ferocious soul meets BFS (Best Face Swap) technology — more natural, and truly unstoppable! 🟦
This is the ultimate speed-optimized Dark Beast V1 evolution, based on Flux.2 Klein 9B,
engineered specifically for lightning-fast low-step + CFG=1 workflows (5steps).
Also available in NVFP4 quantized format, optimized for acceleration on Blackwell architecture GPUs.
( like RTX50XX, PRO6000, B200, and others )
Also supports non-50 series GPUs (automatic 16-bit operation), Verify environment is my ComfyUI 0.11
Key features:
Fully preserves the signature Dark Beast style, rich details, and intense Black Beast
aesthetic from the standard lineage
Refined through advanced targeted distillation & fine-tuning, now perfectly dialed
in for zero-CFG guidance at minimal steps
BlitZ-level inference speed — breathtaking high-quality images in just 5 steps ⚡
Recommended settings: 5 steps, CFG=1 (fixed), any seed you want
In one sentence: Taking Klein’s already blazing speed and cranking it to absolute BlitZ
velocity while keeping every drop of that ferocious Dark Beast soul! 🟦
Lightning-fast generation awaits — unleash it now! 🚀
Usage:
pip install sdnq
py
1import torch
2import diffusers
3from sdnq import SDNQConfig # import sdnq to register it into diffusers and transformers4from sdnq.common import use_torch_compile as triton_is_available
5from sdnq.loader import apply_sdnq_options_to_model
67pipe = diffusers.Flux2KleinPipeline.from_pretrained("GuangyuanSD/FLUX.2-klein-9B-Blitz-Diffusers", torch_dtype=torch.bfloat16)89# Enable INT8 MatMul for AMD, Intel ARC and Nvidia GPUs:10if triton_is_available and(torch.cuda.is_available()or torch.xpu.is_available()):11 pipe.transformer = apply_sdnq_options_to_model(pipe.transformer, use_quantized_matmul=True)12 pipe.text_encoder = apply_sdnq_options_to_model(pipe.text_encoder, use_quantized_matmul=True)13# pipe.transformer = torch.compile(pipe.transformer) # optional for faster speeds1415pipe.enable_model_cpu_offload()1617prompt ="A cat holding a sign that says hello world"18image = pipe(19 prompt=prompt,20 height=1024,21 width=1024,22 guidance_scale=1.0,23 num_inference_steps=4,24 generator=torch.manual_seed(0)25).images[0]2627image.save("flux-klein-Blitz.png")
Original BF16 vs Blitz fine-tune comparison:
Quantization
Model Size
Visualization
Original BF16
18.2 GB
Original BF16
Blitz fine-tune
18.2 GB
DB-Klein2_00005_
Big thanks to @alcaitiff for the awesome work and killer contributions to training Z-Image and Klein models! Seriously impressive stuff! 🚀