TL;DR — Memory-bounded pipeline compresses Qwen3-Coder-30B-A3B (61 GB) to 12.7-16.9 GB (3.6x-4.8x) on a single H200 GPU under a 32 GiB host-RAM ceiling by streaming expert pruning and per-layer RTN W4A16, with WikiText-2 perplexity degradation of +3.9% to +30.5% — no full model ever resident in CPU memory.
ThakiCloud AI Research · 2026-07-30 · 📝 Tech blog (KO)… See the full description on the dataset page:
https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-30-moe-w4a16-pruned-30b-single-gpu.