DeepSeek R1 Distill Qwen 7B Private AI Engine
An optimized, high-performance inference engine using llama.cpp and FastAPI to serve DeepSeek-R1-Distill-Qwen-7B GGUF at ultra-low latency.
🚀 Key Features
- Ultra-Low Latency: Optimized context sizes and thread scheduling tailored for CPU/vCPU containers.
- Reasoning & Coding: Exceptional logic, code explanation, and complete script generation.
- SSE Token Streaming: Sub-50ms first-token response times.
- FIM Autocomplete: Inline completions under 100ms.
- Obsidian Glassmorphic Dashboard: Built-in telemetry dashboard and interactive playground.