Ternary weight (1.58-bit) text-to-image diffusion transformer deployment for NVIDIA GPUs
1.21 GB transformer | 6.4× smaller than FP16 | 4.5 s / 1024² on RTX 3080 | 2.8 s / 1024² on A100 | runs natively on Linux and Windows
Highlights
1.21 GB diffusion transformer, down from 7.75 GB for the FP16 FLUX.2 Klein 4B transformer
Ternary {-1, 0, +1} transformer weights with FP16 group-wise scaling in the matrix-heavy transformer layers (Q/K/V projections, output projections, MLP weights)
Quality-oriented Bonsai Image variant: the additional zero state improves visual quality and prompt fidelity while keeping the transformer compact
4.55 GB CUDA deployment payload including the 4-bit text encoder and FP16 VAE — text encoder is offloaded after prompt encode, so the denoising loop only keeps the compact transformer and VAE resident
4-step FlowMatch-Euler sampler with guidance = 1.0 and shift = 3.0 — no CFG, no negative prompts needed
Gemlite low-bit GEMM path for NVIDIA GPUs, with HQQ used for the compressed text encoder
Runs on Linux and Windows natively through the same CUDA / Gemlite deployment stack
Cross-platform companion: also available as MLX 2-bit for Apple Silicon
Resources
White Paper — full benchmarks, kernels, and memory analysis
Demo repo — one-command setup for Mac / Linux / Windows
This gives an idealized 9.4× reduction relative to FP16 for the ternary transformer layers. A small set of precision-sensitive supporting tensors remains in FP16, so the final Ternary Bonsai Image 4B diffusion transformer is 1.21 GB, a 6.4x reduction from the 7.75 GB FP16 FLUX.2 Klein 4B transformer.
The ternary representation is applied to the matrix-heavy transformer layers, including Q / K / V projections, output projections, MLP linears, and the double-stream add-K / Q / V linears. Supporting tensors (less than 5% of the total parameters) such as modulation streams, embedders, output norm, and output projection remain FP16 for image quality and stability.
The CUDA deployment uses a Gemlite INT2 packed format. Ternary values are stored in 2-bit slots, with the fourth code unused. The model-level Bonsai representation is 1.21 GB; the deployed CUDA pack is 1.54 GB on disk due to runtime packing and alignment overhead in the current Gemlite path.
Memory
Format
Transformer size
Reduction
Ratio
FP16 FLUX.2 Klein 4B
7.75 GB
—
1.0×
Ternary Bonsai Image 4B
1.21 GB
84.4%
6.4×
CUDA deployment:
Component
Size
Gemlite INT2 diffusion transformer
1.54 GB
HQQ 4-bit text encoder
2.84 GB
FP16 VAE
0.17 GB
Total payload
4.55 GB
At runtime, the text encoder is offloaded after prompt encoding. During denoising, the repeated image-generation loop is dominated by the compact ternary diffusion transformer and active image-generation components rather than the full payload.
Peak HBM at 1024² on RTX 3080 is ~6.8 GiB end-to-end (transformer + VAE + activation memory).
Best Practices
Sampler: FlowMatchEuler-discrete with 4 steps, guidance = 1.0, shift = 3.0. The model is designed for 4 steps; running more steps does not improve quality significantly and can introduce artifacts.
Resolution: native 1024² is the design target. 512² works for quick previews.
Aspect ratios: multiples of 32 are supported, including 832x1248 and 1248x832.
Prompting: natural-language prompts. Negative prompts are not required.
Runtime memory: the text encoder is offloaded after prompt encoding, so the denoising loop is memory-light.
Quickstart
Bonsai Studio (Linux / Windows)
The simplest path is the Bonsai Image Demo repo, which sets up the full Bonsai Studio (FastAPI backend + Next.js frontend) and selects gemlite automatically on Linux / Windows:
bash
1git clone https://github.com/PrismML-Eng/Bonsai-Image-Demo.git
2cd Bonsai-Image-Demo
3./setup.sh
4./scripts/download_model.sh # ternary is the default5./scripts/serve.sh
1from backend_gpu.server import build_pipeline
23pipe = build_pipeline(model_id="prism-ml/bonsai-image-ternary-4B-gemlite-2bit")4image = pipe(5 prompt="A bonsai tree in a quiet ceramic studio, soft morning light",6 num_inference_steps=4,7 guidance_scale=1.0,8 height=1024,9 width=1024,10).images[0]11image.save("bonsai.png")
Throughput (CUDA / gemlite)
Warmed wall-clock per image, 4 denoising steps, guidance = 1.0, matched prompts and sampler settings.
Platform
512² (s)
1024² (s)
Notes
A100 (Colab)
1.1
2.8
Ampere datacenter (40 GB)
RTX PRO 6000 Blackwell (Colab)
1.0
2.1
NVIDIA Blackwell, 96 GB VRAM
RTX 3080 10 GB
1.4
4.5
Ampere consumer; 6.8 GiB peak HBM at 1024²
RTX 3060 6 GB (laptop)
3.3
17.5
Ampere mobile; memory-bound at 1024²
The sub-2-bit pack keeps generation viable on commodity GPUs at 1024². The RTX 3080 10 GB reaches 4.5 s/image, while the 6 GB laptop RTX 3060 is the memory-constrained tail.
Benchmarks
Evaluated with matched generation settings across the comparison set on H100. GenEval uses the official 512x512 protocol. For HPSv3 and DPG-Bench, larger-backbone rows are evaluated at 1024x1024, while smaller-backbone rows are evaluated at their native 512x512 setting. Higher is better for all three benchmarks.
Model
Transformer (GB)
GenEval
HPSv3
DPG-Bench
Bonsai Image · Ternary 4B
1.21
0.723
12.22
0.851
Bonsai Image · Binary 4B
0.93
0.671
11.15
0.822
FLUX.2 Klein 4B
7.75
0.819
12.84
0.853
FLUX.1-schnell
23.8
0.716
12.67
0.848
SDXL
5.14
0.300
10.05
0.740
PixArt-Σ XL 2
1.20
0.541
11.93
0.769
Stable Diffusion 1.5
1.72
0.396
4.20
0.601
BK-SDM-Small
0.98
0.297
3.05
0.559
The benchmark results show the intended quality-footprint trade-off. Ternary Bonsai Image 4B is the quality-oriented variant: at 1.21 GB, it sits very close to FLUX.2 Klein 4B across GenEval, HPSv3, and DPG-Bench while reducing the diffusion transformer footprint by 6.4x. The binary companion is the footprint-oriented variant, reducing the diffusion transformer below 1 GB while still delivering strong benchmark results.
Together, the Bonsai Image variants move the quality-footprint frontier: they bring modern diffusion-transformer behavior into a memory range previously occupied by much smaller, lower-capability models.
Use Cases
Local creative tooling: image generation directly on CUDA-equipped workstations and consumer GPUs
Private generation: prompts and generated assets can remain in local or controlled environments
Rapid iteration: lower local latency and no remote queue for iterative creative workflows
Commodity-GPU serving: lower transformer footprint and reduced memory pressure for serving on NVIDIA GPUs
Windows and Linux deployment: native paths through the same Gemlite deployment stack
Enterprise and controlled inference: local or private environments for data residency and compliance-sensitive workflows
Limitations
Ternary Bonsai Image 4B is not bit-identical to the FP16 FLUX.2 Klein 4B model; it is a compact ternary-weight deployment designed to deliver similar practical behavior at much smaller size.
Image-generation quality remains prompt- and workflow-dependent. Small text, fine details, object counts, and strict compositional constraints should be evaluated for the target use case.
Current commodity inference stacks do not yet expose fully native ternary execution as a standard hardware path. This release uses practical Gemlite low-bit GEMM kernels on CUDA.
After the diffusion transformer is made compact, other components such as the VAE can become more visible memory bottlenecks. The runtime mitigates this with text-encoder offload and tiled VAE decoding.
Citation
bibtex
1@techreport{bonsaiimage4b,
2 title = {Bonsai Image 4B: Low-Bit Diffusion on Apple Silicon and Consumer GPUs},
3 author = {Prism ML},
4 year = {2026},
5 month = {May},
6 url = {https://prismml.com}
7}