Views
No views yet
⚠️ EXPERIMENTAL — do not rely on this checkpoint.It was converted before the hybrid-config-drop converter fix: the hybrid (Gated-DeltaNet) portion of the config was dropped during conversion, so the file may not load or run correctly. A corrected reconversion is pending. Even once reconverted, full decode additionally requires the Track-2 DeltaNet int8 kernel (currently a fp16 linear-attention fallback).
Qwen/Qwen3.5-0.8B. Repackaged to the .fni8 resident format (~1.4 GB) for the fni8 W8A8/W4A8 DP4A kernels on NVIDIA Volta (sm_70) — Tesla V100 / CMP 100-210.__dp4a CUDA-core intrinsic. On the CMP 100-210 fleet (whose fp16 tensor cores are firmware-limited) dp4a is the fast path, not a compromise.load_fni8_state_dict(<file>) into an LLMEngine; the architecture is read from the file). ComfyUI-fni8 is for diffusion DiTs only and does not load this model.