This is
Tesslate/OmniCoder-9B (a coding/agentic fine-tune of Qwen3.5-9B) converted to MLC format with
q4f16_1 quantization for in-browser inference via
WebLLM (WebGPU) and
MLC-LLM.
This repo contains the converted weights + mlc-chat-config.json + tokenizer. It can reuse the existing prebuilt Qwen3.5-9B WebGPU model library (same architecture/dimensions), so no separate model-library upload is needed for the default q4f16_1 / 4096-context setup.
I have separtely included in the lib directory recompiled WASM files for 8k, 16k and 32k context windows.