q4f16_1 build of the Beholder state-extractor — runs fully in the browser on
WebGPU via WebLLM, and against any OpenAI-compatible endpoint.version.json and offers a one-click update when a newer build publishes here.params_shard_*.bin + tensor-cache*.json — quantized 4-bit weightsBeholder-q4f16-webgpu.wasm — compiled WebGPU model library (WebLLM model_lib)mlc-chat-config.json — runtime config (browser-right-sized context)tokenizer.json / tokenizer_config.jsonversion.json — update manifest