Views
No views yet
TQ2_0 ternary type (~2.06 bits/weight), using stock llama.cpp (tag b9498) — not PrismML's custom fork/packing.llama-server's OpenAI-compatible endpoint on CPU (Metal has no TQ1_0/TQ2_0 matmul kernels at this llama.cpp version — this model must run with -ngl 0 / CPU-only until upstream adds Metal support):<tool_call> format): pass, clean structured tool_calls output.