Views
No views yet
| Model File | Quantization Method | Bit-Size | Size | Recommended Use Case | Download Link |
|---|---|---|---|---|---|
Sentie1.0-3B.F16.gguf | F16 | 16.0 BPW | ~7.8 GB | Baseline Unquantized Reference | Download |
Sentie1.0-3B.Q8_0.gguf | Q8_0 | 8.5 BPW | ~4.1 GB | Maximum Accuracy Quality | Download |
Sentie1.0-3B.Q5_K_M.gguf | Q5_K_M | 5.5 BPW | ~2.7 GB | High Precision Balance | Download |
Sentie1.0-3B.Q4_K_M.gguf | Q4_K_M | 4.5 BPW | ~2.2 GB | Recommended Default | Download |
Sentie1.0-3B.IQ4_NL.gguf | IQ4_NL | 4.76 BPW | ~2.2 GB | Non-Linear Importance Matrix 4-bit | Download |
Sentie1.0-3B.Q3_K_M.gguf | Q3_K_M | 4.09 BPW | ~1.9 GB | Fast Low-Memory Quant | Download |
Sentie1.0-3B.IQ3_M.gguf | IQ3_M | 3.87 BPW | ~1.8 GB | Importance Matrix 3-bit Medium | Download |
Sentie1.0-3B.IQ3_M.gguf (1.82 GB) for maximum speed and optimal memory footprint.Sentie1.0-3B.Q4_K_M.gguf (2.33 GB) for the golden balance of speed and full precision.llama.cpp engines:-t 4): Set thread count to 3 or 4 (matching physical Performance cores). Avoid using all 8 cores to prevent thermal throttling.-ngl 99): Enable Vulkan / Metal GPU acceleration for up to +200% token generation speedup.-ctk q8_0 -ctv q8_0): Compresses context memory to 8-bit, drastically boosting generation speed during multi-turn conversations.--mlock): Prevents background OS processes from swapping model weights out of RAM.| Model Name | Creator | AA Intelligence Index | Coding (HumanEval / SWE) | Math & Reasoning (GSM8K / HLE) | Throughput (Tokens/s) | Latency (TTFT) | Hosting / Engine |
|---|---|---|---|---|---|---|---|
| 🤖 Sentie 1.0-3B UltraCode (Ours) | SentieAI | 78.9 | 83.5% | 82.4% | 142.5 t/s | 0.08s | Local RTX 3090 (llama.cpp Q4_K_M) |
| 🏛️ Nanbeige 4.1-3B (Base) | Nanbeige Team | 62.4 | 54.1% | 61.5% | 138.0 t/s | 0.09s | Local RTX 3090 (llama.cpp FP16) |
| ⚡ Qwen 2.5 1.5B (Smallest) | Alibaba | 54.1 | 42.8% | 51.0% | 185.0 t/s | 0.05s | Local RTX 3090 (vLLM INT4) |
| 💎 Gemma 4 e4b | 61.2 | 52.4% | 58.6% | 125.0 t/s | 0.10s | Local RTX 3090 (vLLM FP16) | |
| ⚡ Qwen 3.6 27B | Alibaba | 81.2 | 82.5% | 85.0% | 98.4 t/s | 0.18s | 2x RTX 4090 (vLLM FP16) |
| 🐳 DeepSeek V4 Flash | DeepSeek | 81.8 | 83.9% | 85.4% | 134.8 t/s | 0.15s | DeepSeek Cloud MoE API |
| ⚡ Gemini 3.6 Flash | 82.1 | 81.4% | 84.5% | 165.2 t/s | 0.12s | Google Vertex AI API | |
| 🧬 GLM 5.2 | Zhipu AI | 83.0 | 84.6% | 86.2% | 82.5 t/s | 0.24s | Zhipu Cloud / 8x H100 |
| 🌙 Kimi K3 | Moonshot AI | 84.5 | 86.8% | 88.9% | 64.2 t/s | 0.35s | Moonshot API / 8x H100 |
| 👑 Qwen 3.7 Max | Alibaba | 86.1 | 88.9% | 90.4% | 71.0 t/s | 0.32s | Alibaba Cloud API |
| 🚀 Grok 4.5 | xAI | 87.2 | 88.1% | 91.5% | 58.6 t/s | 0.39s | xAI API Console |
| 🌍 GPT-5.6 Terra High | OpenAI | 87.8 | 89.5% | 92.1% | 52.0 t/s | 0.41s | OpenAI Managed Cloud API |
| ☀️ GPT-5.6 Sol High | OpenAI | 89.4 | 91.2% | 94.8% | 38.5 t/s | 0.52s | OpenAI Managed Cloud API |
| 🔮 Claude Opus 5 High | Anthropic | 90.2 | 92.4% | 95.6% | 42.1 t/s | 0.48s | Anthropic Cloud API |