Views
No views yet
beta2_pruned) — the MiniMax-H3 graft experiment carrying character from LTX-2.3 / Wan 2.2 / Krea 2.
The author ships BF16 only (40GB); this is the same model at less than a third of the size..comfy_quant), quantize exactly those, and pass
the other 332 tensors through untouched. Key structure verified against the official file:
0 missing / 0 extra / 0 shape mismatches. Output lands at the same 12.53GB.attn_k warning ("touching K silently degrades audio"): that concern is about
weight surgery. For quantization, the official H3 NVFP4 already quantizes the fused qkv_proj
in production — this repo follows that proven layout. Bake script included
(bake_eros_nvfp4.py), ~24s on one RTX PRO 2000 Blackwell.models/diffusion_models/, load with UNETLoader, and reuse your existing
MiniMax-H3 workflow as-is: CLIPLoader type minimax, H3 video/audio VAE,
MiniMaxH3ImageToVideo, MiniMaxH3SigmaShift (12.0 / 3.0), res_multistep.