Views
No views yet
Oobabooga, llama.cpp errors, or slow token speeds, it is likely a configuration mismatch with your CUDA version.
The Easiest Way to Run This Model:| GPU | VRAM | Recommended Model | Status |
|---|---|---|---|
| RTX 4090 | 24GB | DeepSeek R1 (Q4_K_M) | ✅ Verified (V6rge) |
| RTX 3060 | 12GB | DeepSeek R1 (Q2_K) | ✅ Verified (V6rge) |
| Mac M1/M2/M3 | Shares RAM | DeepSeek R1 (Q4) | ✅ Verified (V6rge) |
flash-attention by default if supported.