Qwen text encoder split into 9 UINT8 .sentis stages
FLUX.2 transformer split into 7 FP16 .sentis stages
VAE decoder split into 6 FP16 .sentis stages
Qwen tokenizer JSON used by the Unity C# tokenizer
A SHA-256 manifest for all runtime files
The 23 runtime files total approximately 10.23 GiB. Staging the networks keeps
peak memory substantially below loading the complete text encoder,
transformer, and VAE at the same time.
Verified Unity result
This exact uploaded artifact set was tested end-to-end with:
Unity 6000.5.5f1
com.unity.ai.inference2.6.1
Windows Editor using Direct3D 12
NVIDIA GeForce RTX 4090 with 24 GB VRAM
LocalImageGen's accelerated D3D12 path: the seven FP16 transformer stages
run through persistent GPUCompute workers, while the UINT8 text encoder
and FP16 VAE run on CPU to keep the complete pipeline within 24 GB VRAM
1024 x 1024 output
4 inference steps
Guidance scale 1.0
Seed 27101
Prompt:
A copper clockwork fox carries 27 blueberries through Sao Paulo's moonlit
greenhouse; watercolor, ultra-detailed.
The Unity C# tokenizer produced 38 prompt tokens. Initialization took
2.459 s; generation took 169.656 s; the complete validation took
172.135 s (2m52s). It produced a valid prompt-aligned image, reported no
non-finite tensors, and left the Unity console with zero errors. Peak observed
GPU memory was approximately 22.6 GB.
Verified LocalImageGen output
The tested model-content revision is
fd953005ff22e26b58fce7da6d9a8b6786c97865. LocalImageGen pins this immutable
revision so later model-card-only commits cannot change downloaded runtime
artifacts.
Direct3D 12 support
D3D12 acceleration is supported by the current LocalImageGen runtime. The
important implementation detail is worker lifetime: all seven transformer
workers are created once and reused for every denoising step. Recreating the
seven large workers on every step caused Unity's D3D12 backend to retain heap
reservations and eventually exhaust GPU memory.
The mixed execution path is deliberate:
Transformer: D3D12 GPUCompute
Text encoder: CPU
VAE decoder: CPU
Keeping the text encoder and VAE off the GPU leaves enough room for the
persistent FP16 transformer worker set. The verified RTX 4090 run nearly
filled 24 GB VRAM, so 24 GB should be treated as the recommended D3D12
configuration for this 1024 x 1024 profile. Other GPU and VRAM configurations
have not yet been validated.
LocalImageGen also supports decoding intermediate previews when Show
Intermediates is enabled. Intermediate VAE decoding runs on CPU and therefore
adds noticeable time; early denoising previews are expected to look noisy.
This remains the high-quality local demonstration tier rather than the
lightweight desktop or mobile option.
This is a mixed UINT8/FP16 build. The text encoder uses UINT8 weight storage;
the transformer and VAE use FP16 to preserve generation and decoded image
quality.
A naively UINT8-quantized transformer was tested and rejected: it retained
coarse prompt semantics but produced severely washed, yellow-tinted images.
Those transformer artifacts are intentionally not included in this release.
Using it
Install Local Image Generator in a Unity 6 project, then use:
Tools > LocalImageGen > Model Setup
Select Modern 1024 — FLUX.2 Klein 4B, download the pinned repository
revision, and assign the downloaded files to the FLUX.2 generator prefab.
Inference and tokenization then run inside Unity through the package's public
C# API.
Conversion notes
Unity Inference Engine's ONNX importer does not preserve arbitrary offsets in
external tensor data. The source graphs were therefore repacked so external
initializers are sequential and offset-bearing LayerNorm constants are
embedded. Incorrect symbolic rotary-batch metadata was also repaired before
the .sentis files were generated.
Python-based tooling was used to convert and validate the source model and to
publish this repository. Python is not part of the Unity runtime or public
package API.
Provenance and license
The source model is
black-forest-labs/FLUX.2-klein-4B.
The source model and this converted artifact set are distributed under the
Apache License 2.0; see LICENSE.
FLUX is a model family from Black Forest Labs. This repository is an
independent Unity conversion and is not an official Black Forest Labs or
Unity Technologies release.