Views
No views yet
coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta, 2026-06)..aimodel)[!NOTE] Update 2026-07-15:gpu-pipelined-b2/addsqwen3_5_2b_decode_int8hu_block32_symre-exported withcoreai-core 1.0.0b2, loadable on the OS 27 beta 3 toolchain (June-era b1 bundles fail to load there with a versioned-IR error). The original b1 tree is retained unchanged so existing apps and pinned catalogs keep working. The b2 decode bundle is the exact artifact measured on DeviceMark.
coreai-pipelined
GPU engine via the decode-only loop-free export — async encode, on-GPU argmax
sampling, on-device KV growth, zero custom kernels.| surface (ship bundle) | prefill (S=1) | decode |
|---|---|---|
M4 Max (release llm-benchmark, p=128 g=256) | 161.2 | 160.8 tok/s |
| iPhone 17 Pro (one-shot runner, 2 runs × 2 trials) | 29.7–30.3 | 28–30 tok/s — ≥ the CoreML qwen3.5-2B port (~27) |
import CoreAIOps; no session, no model plumbing, downloads on first use):let tldr = try await CoreAI.summarize(text, options: .model("qwen3.5-2b"))1git clone https://github.com/john-rocky/coreai-kit
2open coreai-kit/Examples/ChatDemo/ChatDemo.xcodeproj
3# → Run, then pick "Qwen3.5 2B" in the model picker
4
5# agents / headless (macOS):
6cd coreai-kit/Examples/ChatDemo
7swift run chat-cli --model qwen3.5-2b --prompt "What can you do, offline?"1import CoreAIKit
2
3let chat = try await ChatSession(catalog: "qwen3.5-2b")
4let reply = try await chat.respond(to: prompt)
5// reply: the answer, generated fully on-deviceKitLanguageModel plugs this bundle into the system LanguageModelSession; capabilities (tool calling, guided generation) auto-detect per model.Examples/ChatDemo/Sources/QuickStart.swift
— this exact code as one typed function, no UI; the CLI is an argument shell over it, and
the GUI drives the same ChatSession across turns for its transcript.
Multi-turn? Hold the ChatSession and call respond(to:) per turn — it keeps the
conversation history; streamResponse(to:) yields tokens as they decode.https://github.com/john-rocky/coreai-kit → product CoreAIKitcom.apple.developer.kernel.increased-memory-limitdownloadProgress callback)gpu-pipelined/qwen3_5_2b_decode_int8hu_perchan_sym/ — the ship config (2.9 GB):
transformer int8 linear per-block-32 + untied lm_head in per-block-32 absmax int8
(int8hu --head-sym). The head trick is what unlocks the speed: the
248 K-vocab fp16 head was ~1.0 GB of the ~2.4 GB per-token read. Crucial detail: the head
must be quantized with plain absmax symmetric — the default
symmetric_with_clipping clips outlier head rows and flips oracle top-1s (full story in
the zoo's pipelined-engine notes).
Naming note (2026-06-11): the directory says _perchan_sym, but its head is
per-block-32 — the export script of the day parsed the granularity flag without applying
it (since fixed); byte-identical bundle sizes confirmed it. All numbers were measured on
exactly these bytes and stand. Genuinely per-channel (axis-0) int8 weights are broken
on the current beta GPU delegate (garbage logits), so per-block-32 + symmetric IS the
correct ship shape. The dir name is kept to avoid breaking download paths.gpu-pipelined/qwen3_5_2b_decode_int8lin/ — fp16-head variant (2.4 GB): 127 tok/s Mac /
19–21 iPhone. Smaller; keep if you want the head at full precision.metadata.json + tokenizer/ + .aimodel), input_ids
STATIC [1,1] (loop-free single-step GDN), position_ids + KV seq dynamic → EngineFactory
classifies them dynamic → pipelined engine.apps/coreai-shared-product.patch →
apps/coreai-pipelined-extra-states.patch; Apple's repo is issues-only, so capabilities ship
as patches), then:COREAI_CHUNK_THRESHOLD=1 llm-benchmark --model qwen3_5_2b_decode_int8hu_perchan_sym -p 128 -g 256 -n 3COREAI_CHUNK_THRESHOLD=1 before engine creation — prefill runs as pipelined S=1 steps
(prompt tok/s ≈ decode tok/s).engine.warmup() — it warms query length 256 and the static [1,1] graph
rejects it. A 1-token generate after load is the warmup (llm-runner needs
--warmup exact --warmup-length 1).com.apple.developer.kernel.increased-memory-limit entitlement — cold GPU
specialization dies with std::bad_alloc at the default jetsam limit without it.NSPOSIXErrorDomain code=2 at engine create — uninstall the app to
reclaim.conversion/export_qwen3_5_decode_pipelined.py
(int8hu --head-sym --hf-id Qwen/Qwen3.5-2B) ·
knowledge/pipelined-engine.md