These models are made to work with
stable-diffusion.cpp release
master-ac54e00 onwards. Support for other inference backends is not guarenteed.
Normal K-quants are not working properly with SD3.5-Large models because around 90% of the weights are in tensors whose shape doesn't match the 256 superblock size of K-quants and therefore can't be quantized this way. Mixing quantization types allows us to take adventage of the better fidelity of k-quants to some extent while keeping the model file size relatively small.
Only the second layers of both MLPs in each MMDiT block of SD3.5 Large models have the correct shape to be compatible with k-quants. That still makes up for about 10% of all the parameters.
Sorted by model size (Note that q4_0, q4_k_4_0, and iq4_nl are the exact same size)
Generated with a modified version of sdcpp with
this PR applied to enable clip timestep embeddings support.
Text encoders used: q4_k quant of t5xxl, full precision clip_g, and q8 quant of
ViT-L-14-TEXT-detail-improved-hiT-GmP-TE-only-HF in place of clip_l.
Full prompts and settings in png metadata.