Views
No views yet
Qwen/Qwen3-8B. Repackaged to the .fni8 resident format (~8.8 GB) for the fni8 W8A8/W4A8 DP4A kernels on NVIDIA Volta (sm_70) — Tesla V100 / CMP 100-210.models/COVERAGE.md). Performance is fleet-specific. All fni8 speedups are measured on the CMP 100-210 mining-card fleet, where the fp16 tensor cores are firmware-gimped. These numbers do not transfer to a real Tesla V100 (whose fp16 tensor cores would beat dp4a).__dp4a CUDA-core intrinsic. On the CMP 100-210 fleet (whose fp16 tensor cores are firmware-limited) dp4a is the fast path, not a compromise.load_fni8_state_dict(<file>) into an LLMEngine; the architecture is read from the file). ComfyUI-fni8 is for diffusion DiTs only and does not load this model.