Views
No views yet
| Model | Mean PSNR (dB) | Std (dB) | Median (dB) | P5 (dB) | P95 (dB) |
|---|---|---|---|---|---|
| dinac_ae_d2 | 35.59 | 4.87 | 35.40 | 27.89 | 43.51 |
| dinac_ae | 35.19 | 4.53 | 35.06 | 28.02 | 42.43 |
| FLUX.2 VAE | 36.28 | 4.53 | 36.07 | 28.89 | 43.63 |
2000 validation images as DINAC-AE. FLUX.2 numbers are
reused from the existing DINAC-AE 2k benchmark and were not recomputed for this
export.35.46 dB mean PSNR (25.61 min, 46.69 max).NVIDIA GeForce RTX 5090 in bfloat16, averaging repeated
batches per resolution.| Resolution | Batch Size | Model | Encode (ms/batch) | ms/image | Images/s | Peak VRAM (MiB) | Speedup vs FLUX.2 | Peak VRAM Reduction vs FLUX.2 |
|---|---|---|---|---|---|---|---|---|
256x256 | 128 | dinac_ae_d2 | 69.56 | 0.543 | 1840.0 | 1606.5 | 4.92x | 87.2% |
256x256 | 128 | dinac_ae | 50.25 | 0.393 | 2547.4 | 1569.7 | 6.80x | 87.5% |
256x256 | 128 | FLUX.2 VAE | 341.94 | 2.671 | 374.3 | 12533.8 | 1.00x | 0.0% |
512x512 | 32 | dinac_ae_d2 | 75.09 | 2.347 | 426.2 | 1606.7 | 4.74x | 87.2% |
512x512 | 32 | dinac_ae | 53.09 | 1.659 | 602.7 | 1570.0 | 6.70x | 87.5% |
512x512 | 32 | FLUX.2 VAE | 355.64 | 11.114 | 90.0 | 12533.8 | 1.00x | 0.0% |
encode() returns DINAC-AE-D2's own whitened latent space.decode() expects that same whitened latent space and dewhitens internally.predict_class() expects the same whitened latent space, dewhitens
internally, and predicts a DINOv2-B class-token feature.whiten() and dewhiten() are exposed for explicit control.encode_posterior() returns the raw exported posterior before whitening.DinacAEInferenceConfig.num_steps counts decoder evaluations directly:
num_steps=1 means one NFE.float32. The recommended runtime path is
bfloat16 AMP for the main encoder, decoder, and class-token path. The loader
retains normalization affine parameters, GRN/residual gates, the final pixel
projection, latent statistics, RoPE/time frequencies, sampler state, and
whitening/dewhitening in float32. These tensors are loaded from the original
FP32 weights before ordinary parameters are converted to BF16.1import torch
2
3from dinac_ae import DinacAE, DinacAEInferenceConfig
4
5
6device = "cuda"
7model = DinacAE.from_pretrained(
8 "data-archetype/dinac_ae_d2",
9 device=device,
10 dtype=torch.bfloat16,
11)
12
13image = ... # [1, 3, H, W] in [-1, 1], H and W divisible by 16
14
15with torch.inference_mode():
16 latents = model.encode(image.to(device=device, dtype=torch.bfloat16))
17 class_token = model.predict_class(latents)
18 recon = model.decode(
19 latents,
20 height=int(image.shape[-2]),
21 width=int(image.shape[-1]),
22 inference_config=DinacAEInferenceConfig(num_steps=1),
23 )8-block ViT/DiT-style transformer encoder and an
8-block FCDM decoder.16, model width is 896, and latent width is 128.154.22M: 78.02M encoder, 61.93M decoder, and
14.26M DINO token/class alignment head.predict_class(latents) exposes the DINOv2 ViT-B/14 class-token feature
directly from latents.1@misc{dinac_ae_d2,
2 title = {DINAC-AE-D2: a DINOv2-aligned class-token diffusion autoencoder},
3 author = {data-archetype},
4 email = {data-archetype@proton.me},
5 year = {2026},
6 month = jun,
7 url = {https://huggingface.co/data-archetype/dinac_ae_d2},
8}