⚠️⚠️⚠️ This is at the moment a debugging upload. It will trigger a vllm assert as some scaling factors are not >0.0 if used with FP8 KV-cache (--kv-cache-dtype='fp8')
📥 Usage & Running Instructions
The model was tested with vLLM and 2x RTX Pro 6000, here is a script suitable for such configuration.