Views
No views yet
coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta, 2026-06)..aimodel bundles from Apple's official
coreai-models export recipe — unmodified,
with the exact environment, hashes, and measured performance published.1uv run coreai.llm.export qwen3-0.6b # macOS
2uv run coreai.llm.export qwen3-0.6b --platform iOSimport CoreAIOps; no session, no model plumbing, downloads on first use):let tldr = try await CoreAI.summarize(text, options: .model("qwen3-0.6b"))1git clone https://github.com/john-rocky/coreai-kit
2open coreai-kit/Examples/ChatDemo/ChatDemo.xcodeproj
3# → Run, then pick "Qwen3 0.6B" in the model picker
4
5# agents / headless (macOS):
6cd coreai-kit/Examples/ChatDemo
7swift run chat-cli --model qwen3-0.6b --prompt "What can you do, offline?"1import CoreAIKit
2
3let chat = try await ChatSession(catalog: "qwen3-0.6b")
4let reply = try await chat.respond(to: prompt)
5// reply: the answer, generated fully on-deviceExamples/ChatDemo/Sources/QuickStart.swift
— this exact code as one typed function, no UI; the CLI is an argument shell over it, and
the GUI drives the same ChatSession across turns for its transcript.
Multi-turn? Hold the ChatSession and call respond(to:) per turn — it keeps the
conversation history; streamResponse(to:) yields tokens as they decode.https://github.com/john-rocky/coreai-kit → product CoreAIKitdownloadProgress callback).aimodel is a build artifact, not a pure function of the recipe — the
same export command produced a 2.2× slower artifact across the macOS 26 → 27β
boundary (forensics).
Hosted artifacts + hashes are the reproducible ground truth; every bundle here
is exactly the one measured in
apple-silicon-llm-bench.| Bundle | Contents | SHA-256 (main.mlirb) |
|---|---|---|
macos/ | macOS dynamic, int4 (macOS-27β export) | e05ad9093c651e07e0a9c8589319ec1cc9e865e2b474f52a22b143d9ab6c3147 |
macos-26-export/ | macOS dynamic, int4 — macOS-26-era artifact, 2.2× faster, cannot be re-created on 27β | f7a8357f50292f4425591fb0ed2ef4c89c91b658498d89e7e8b516eca0e89554 |
ios/ | iOS static ctx4096, mixed 4/8-bit palettized | 151bbb15ef14b599bc62b7b08c2969e732febe0a2d43886414aa9d5f29213b01 |
llm-benchmark, greedy)| Bundle | Protocol | Decode tok/s | Prefill | Load (warm) |
|---|---|---|---|---|
| macos (27β) | M4 Max, 512p/1024g | 484 | 9,396 | 0.10 s |
| macos-26-export | M4 Max, 512p/512g warm | 1,121 | — | — |
| macos-26-export | iPhone 17 Pro GPU (h18p), 512p/1024g | 115 | 5,807 | 0.07 s |
| ios (ANE, h18p) | iPhone 17 Pro, 512p/1024g | 69.6 | 5,325 | 0.045 s |
coreai-core 1.0.0b1 · coreai-torch 0.4.0 · coreai-opt 0.2.0 · torch 2.9.0b1cb71b (export code identical to upstream 0c1055f)1# CLI (from a coreai-models checkout)
2swift run -c release llm-runner --model <downloaded-bundle-dir> --prompt "Hello"
3swift run -c release llm-benchmark --model <downloaded-bundle-dir>xcrun coreai-build compile <ir>.aimodel --platform iOS --preferred-compute neural-engine --architecture h18p
(h18p = iPhone 17 Pro), then set metadata.json assets.main to the .aimodelc.