Place in model_splits/ (no splitting needed — single file)
node serve.js (port 8180)
Open http://localhost:8180 in Chrome
Use Cases
Lightweight chat and Q&A
Classification and summarization
Edge/IoT inference
Testing and prototyping
Hardware
Any WebGPU-capable device. Tested on AMD Strix Halo but works on much smaller hardware too. The model is only 369 MB — it fits anywhere.
Why This Package
Part of a series making popular models available on WebGPU for AMD unified memory AI PCs. WebGPU bypasses broken ROCm and routes through the gaming driver stack.