Views
No views yet
dtype: "q4" with AutoModelForCausalLM.from_pretrained. This artifact passed ONNX CPU inference and a pinned Transformers.js 4.2.0 WebGPU generation smoke. Representative devices, Worker mode, multi-turn and production-length RAG, browser memory, and quantized heldout quality still require qualification before migration.df58c174f05ff733f83f8cae10ea9298224c8006.Kanha-AI/kanha-liquid-lfm2.5-1.2b-instruct-sft-v1-adapter@50869bbad65a6b9222ccdbc2cfe332541fc178f6.Liquid4All/onnx-export@9a23ddd23035165f7414a5de3220a51e85780f64.chat_template.jinja removes only the training-time generation and endgeneration Jinja markers, which Transformers.js 4.2.0 does not parse. Rendered single-turn and multi-turn inference prompts were verified byte-for-byte against the native template before publication.tokenizer_config.json, which is the source Transformers.js loads.