A frontier coding & agent model that runs on your desk. Q4_K_M / IQ3_XXS GGUF of tencent/Hy3 (Hunyuan 3.0 — 295B total, 21B active). Quantized directly from official Tencent BF16 weights by BatiAI — code+multilingual‑calibrated imatrix, MTP‑pruned, BatiAI‑signed.
| Hy3 | GLM‑5.2 | DeepSeek‑V4 | |
|---|---|---|---|
| Total params | 295B | 753B | ~1.6T |
| Active / token | 21B | 40B | ~37B |
| Fits a 128GB Mac? | ✅ (IQ3_XXS) | ✗ | ✗ |
| SWE‑Bench Verified | SWE‑Bench Pro | GPQA Diamond | BrowseComp |
|---|---|---|---|
| 78.0 | 57.9 | 90.4 | 84.2 |
| Quant | Size | Min RAM | Best for | Quality |
|---|---|---|---|---|
| Q4_K_M | 166 GB (4 shards) | 192 GB | 256GB Mac Studio / server | ⭐ Cleanest — recommended when RAM allows |
| IQ3_XXS | 106 GB (3 shards) | 128 GB | 128GB Mac Studio | ✅ Great — fits a 128GB Mac (raise the Metal wired limit; ~106 GiB of weights leaves modest context room) |
--prune-layers 80) — the speculative head gives
no benefit on Apple Metal and isn't imatrix‑covered, so a clean 80‑layer text model is the right target.binary_search (verify log hy3-q4-verify.log shipped in this repo):1# prompt: def binary_search(arr, target):
2 lo, hi = 0, len(arr) - 1
3 while lo <= hi:
4 mid = (lo + hi) // 2
5 if arr[mid] == target:
6 return mid
7 elif arr[mid] < target: lo = mid + 1
8 else: hi = mid - 1
9 return -1# 测试) — exactly the low‑bit zh mixing flagged below. Logic was correct; use Q4_K_M for the cleanest output.⚠️ Positioning — read this. Hy3's strength is frontier coding / reasoning / agentic tool‑calling (EN/ZH). It is not a Korean‑specialized model (Tencent origin, no published Korean benchmark); lower‑bit quants can show occasional zh/en token mixing on Korean — use Q4_K_M for the cleanest Korean. For Korean‑first chat/STT on 16GB Macs, use batiai/qwen3.6‑27b. Hy3 is a frontier / high‑RAM tier model (like Kimi K2.6, GLM‑5.1, DeepSeek‑V4) — 128GB+ Apple Silicon or a workstation/server only.
⚙️ Build: Hy3 (hy_v3arch) needs hy_v3 support — mainline merge pending (ggml‑org/llama.cpp#25395); build from that PR for now. Ollama support follows the mainline merge.⚠️ Chat template: the stock Hy3 Jinja template uses.format()calls llama.cpp rejects. This repo ships a fixed template (Hy3-chat_template.jinja) — pass it with--jinja.
1# 1) download — sharded GGUF (llama.cpp auto‑loads all shards from the first one)
2# 128GB Mac → IQ3_XXS | 256GB / server → Q4_K_M
3hf download batiai/Hy3-GGUF \
4 "Hy3-IQ3_XXS-*.gguf" Hy3-chat_template.jinja --local-dir ./hy3
5
6# 2) chat (Apple Silicon Metal)
7./llama-cli -m ./hy3/Hy3-IQ3_XXS-00001-of-00003.gguf -ngl 99 -c 8192 \
8 --jinja --chat-template-file ./hy3/Hy3-chat_template.jinja \
9 -p "Refactor this function and explain the change."
10
11# raw completion (no chat template): add -no-cnvHy3-imatrix.dat (the calibration matrix used) is included for transparency / re‑quantization.1@misc{batiai-hy3-gguf-2026,
2 title = {Hy3 (Hunyuan 3.0) GGUF — code+multilingual calibrated quantization},
3 author = {BatiAI},
4 year = {2026},
5 publisher = {Hugging Face},
6 url = {https://huggingface.co/batiai/Hy3-GGUF}
7}