This is the BF16 conversion of HiDream-O1-Image for use with ComfyUI. Weights have been cast to bfloat16 for a balance of precision and memory efficiency.
A GPU with at least 20 GB VRAM is recommended for comfortable use at full 2048 × 2048 resolution. 24 GB cards (RTX 3090/4090, A5000, etc.) will have no issues.
Open ComfyUI and use the workflow provided in the custom node repository. Point the model loader to HiDream-O1-Image-bf16.
About HiDream-O1-Image
HiDream-O1-Image is a natively unified image generative foundation model built on a Pixel-level Unified Transformer (UiT) — no external VAEs, no disjoint text encoders. It encodes raw pixels, text, and task-specific conditions in a single shared token space, supporting:
At only 9B parameters it matches or exceeds much larger open-source DiTs and leading closed-source models. It debuted at #8 in the Artificial Analysis Text to Image Arena (2026-05-05).
Key Features
🧬 Pixel-Level Unified Transformer — end-to-end on raw pixels, no VAE, no disjoint text encoder
🎨 One Model, Many Tasks — T2I, editing, personalization, storyboard generation
🧠 Reasoning-Driven Prompt Agent — built-in "thinking" agent that resolves layout and rendering before generation
🖼️ Native High Resolution — direct synthesis up to 2,048 × 2,048
⚡ 9B Parameters — performance parity with models many times larger
GenEval (compositional generation) — HiDream-O1-Image scores 0.90 overall at 9B params, second only to the 200B+ Pro variant and ahead of GPT Image 2 (0.89).
DPG-Bench (dense prompt alignment) — Overall score 89.83, ranking second behind the Pro variant.