Views
No views yet
Qwen/Qwen3-VL-4B-Instruct
for the Piccolo AI OpenVINO Model Server provider.ebb281ec70b05090aa6165b016eac8ec08e71b17v2026.2, commit
9b795c5fad08cc4abf06a8f751b80a5fc5ae1001d4dd21a3aa89c0671d85b704847ac06a378e761c2026.2.0rc22026.2.0.0rc25.0.03.2.01.0128INT4_SYM0.1.3
and OVMS 2026.2 on an Intel GPU target:ScaledDotProductAttention operations;ScaledDotProductAttention operations;AVAILABLE;5.65 tok/s at concurrency
1 to 11.52 tok/s at concurrency 8 for a fixed 128-token completion test;8.02 GiB, with 0 max-limit events,
0 OOM events, and 0 swap use.8.005 GiB soft memory boundary./models/model. The Piccolo AI
provider exposes it through the OpenAI-compatible /v3 API under the stable
model identifier piccolo-chat.