Views
No views yet
Qwen/Qwen-Image. Repackaged to the .fni8 resident format (~31.4 GB (int4 DiT)) for the fni8 W8A8/W4A8 DP4A kernels on NVIDIA Volta (sm_70) — Tesla V100 / CMP 100-210.__dp4a CUDA-core intrinsic. On the CMP 100-210 fleet (whose fp16 tensor cores are firmware-limited) dp4a is the fast path, not a compromise.UnetLoaderFNI8 runs the diffusion transformer through the dp4a kernels; the text encoder and VAE are unchanged). fni8-serve is for LLMs only and does not load this model.