Serving recipes, launch wrappers, and the measured throughput/context ladder for
running a W4A16 DeepSeek-V4-Flash-class model on 8x RTX 3090 (SM 8.6, 24 GiB each)
with CUDA graph decode, FlashInfer sparse MLA, Marlin MoE, and compressed hybrid KV.
Concurrent sequences amortize the TP8 allreduce that dominates each decode step,
so aggregate throughput scales near-linear while… See the full description on the dataset page:
https://huggingface.co/datasets/Relativ3pa1n/dsv4-flash-sm86-8x3090.