Qwen2 with 1.5B ONNX INT8 — Mobile-Optimized Quantized Model
📌 Description
This is a quantized INT8 version of the Qwen language model,
converted to ONNX format for efficient inference on low-resource
devices such as mobile phones and edge hardware.
🎯 Objective
The goal of this project is to make Qwen accessible on devices
with limited memory and compute power (e.g. iPhone 12, mid-range
Android phones) without requiring high-end GPUs or cloud infrastructure.