This is an 8-bit quantised CoreML conversion of the
oceanicity/Qwen3-4B-Instruct-2507 model. It has been heavily optimised for fast, efficient, and low-memory inference on Apple Silicon using the Apple Neural Engine (ANE).
Conversion and quantisation work was performed using a customised version of
0seba's coremlmodels tool.
This model is ready to be used in CoreML inference pipelines that support multi-chunked stateful transformers. Ensure that your inference engine stitches the chunks together sequentially and routes the KV cache states appropriately.