Views
No views yet
Intel/Qwen3.5-122B-A10B-int4-AutoRound.Qwen/Qwen3.5-122B-A10B-FP8.--profile dense). It requires that project's vLLM 0.23 hybrid-FP8 dispatch
patch — an INCConfig.maybe_update_config override that dispatches
Fp8LinearMethod for the FP8 shared-expert layers. Stock vLLM will not dispatch
the mixed INT4/FP8 scheme.albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4.
Built with build-hybrid-checkpoint.py, which merges FP8 non-expert tensors into
the INT4 checkpoint.