Views
No views yet
tokenizer.json)Qwen/Qwen-Image ships its Qwen2 BPE tokenizer as vocab.json + merges.txt only — there is
no fast tokenizer.json in the upstream repo (the Python fork builds the fast tokenizer at
runtime via transformers). The Rust engine's tokenizer loader (mlx_gen::TextTokenizer, consumed by
the qwen-image provider's load_tokenizer) reads the HF tokenizers fast serialization, so it
needs a tokenizer.json.tokenizer.json so SceneWorks model-install can overlay it onto the
upstream Qwen-Image snapshot (instead of running a Python vocab.json+merges.txt→fast conversion at
install time on every machine — the desktop Mac bundle ships no Python). See SceneWorks sc-6570; this
mirrors the Kolors fast-tokenizer overlay
(sc-4764).Note:Qwen/Qwen-Image-Edit-2511already ships its owntokenizer.jsonupstream, so only the base text-to-imageQwen/Qwen-Imagerepo needs this overlay.
tools/build_qwen_tokenizer.py (mlx-gen): loads the Qwen2 tokenizer with
transformers.AutoTokenizer.from_pretrained (the fast path) and writes backend_tokenizer.save(...).
The result is the byte-identical fast tokenizer the fork builds at runtime — same vocab, merges,
NFC + ByteLevel pipeline, and special tokens.transformers tokenizer across an
EN + EN-long + CN + mixed CN/EN/numeric/punct + empty(negative-prompt) battery — 0 mismatches.
vocab_size 151665, pad token id 151643 (<|endoftext|>).tokenizer.json — the derived fast tokenizer (the file the Rust engine needs).vocab.json, merges.txt, tokenizer_config.json, added_tokens.json, special_tokens_map.json —
the upstream slow-tokenizer source files (provenance / reproducibility).