VOSR 1.4B Mobile
AllGPTORG/VOSR_1.4B_Mobile is a mobile deployment package of the
VOSR-1.4B one-step image super-resolution model, quantized and compiled as
Qualcomm QNN DLC graphs for the Snapdragon 8 Gen 3 NPU.
The package is designed for high-quality image restoration and super-resolution,
including text-rich images. It preserves the original one-step VOSR pipeline while
splitting the network into smaller graphs suitable for mobile integration.
This repository contains QNN deployment artifacts. It is not a Transformers or
Diffusers checkpoint and cannot be loaded with from_pretrained().
Model summary
| Property | Value |
|---|
| Base model | CSWRY/VOSR, VOSR-1.4B one-step |
| Task | Generative image restoration and super-resolution |
| Target SoC | Qualcomm Snapdragon 8 Gen 3 / SM8650 |
| Compile target | Samsung Galaxy S24 family, Android 14 |
| Runtime format | Qualcomm QNN DLC |
| Static tile size | 512 x 512 pixels |
| Batch size | 1 |
| Activations | FP16 (A16) |
| DiT block weights | INT4 (W4A16) |
| Auxiliary/final graph weights | INT8 (W8A16), except the FP16 VAE decoder |
| Total DLC size | 1,736,469,708 bytes / 1.617 GiB |
Files
The graphs must be executed in the order shown below.
| Order | File | Precision | Size |
|---|
| 1 | vosr_dinov2l_layer17.dlc | W8A16 | 229.50 MiB |
| 2 | vosr_qwen_vae_encoder.dlc | W8A16 | 19.09 MiB |
| 3 | vosr_dit_prepare.dlc | W8A16 | 42.61 MiB |
| 4 | vosr_dit_blocks_00_12.dlc | W4A16 | 444.86 MiB |
| 5 | vosr_dit_blocks_12_18.dlc | W4A16 | 222.62 MiB |
| 6 | vosr_dit_blocks_18_24.dlc | W4A16 | 222.62 MiB |
| 7 | vosr_dit_blocks_24_36.dlc | W4A16 | 444.86 MiB |
| 8 | vosr_dit_final.dlc | W8A16 | 4.90 MiB |
| 9 | vosr_qwen_vae_decoder.dlc | Legacy W8A16 DLC; EPContext uses FP16 | 24.97 MiB |
manifest.json contains the same graph order, precision assignment, and exact
byte size for programmatic use.
Ready-to-run SM8650 EPContext package
sd8g3/qnn-context/ contains eleven pre-linked QNN context binaries for the
Snapdragon 8 Gen 3 HTP. They were linked with QAIRT 2.45.0.260326154327,
target DSP v75 / SoC model 57, and use O1 graph finalization to stay inside the
device's 8 MiB VTCM budget.
The DiT is split into six contiguous stages: 00_06, 06_12, 12_18,
18_24, 24_30, and 30_36. RMSNorm pointwise multiplications are divided
along the token axis before linking; a CPU reference comparison against the
unmodified QDQ block measured 72.9 dB PSNR.
The complete context package is 1,140,031,488 bytes (1.062 GiB). The small
EPContext ONNX wrappers and their exact SHA-256, byte-size, tensor, graph-order,
and immutable Hub-revision contracts are stored in
sd8g3/vosr_runtime_manifest.json.
The Qwen VAE decoder is compiled directly from the original FP16 ONNX graph.
The earlier W8A16 QDQ decoder removed nearly all spatial detail. The FP16
SM8650 context matches the CPU ONNX decoder at 59.77 dB PSNR with a maximum
per-channel error of one RGB level.
Raspberry Pi 5 ONNX Runtime package
pi/w8a16/ contains the CPU-oriented VOSR-1.4B runtime used by AllCamera Cloud
AI. Its DINO and DiT MatMul weights use ONNX Runtime MatMulNBits INT8
quantization while activations remain FP16. VAE, prepare, and final graphs stay
FP16. The package is 1.663 GiB, down from 3.116 GiB for the split FP16 graphs,
and is pinned by the immutable Hub tag pi-w8a16-v1.
The graphs are deliberately executed one at a time so an 8 GB Raspberry Pi 5
does not hold the complete 1.4B pipeline in memory. The largest measured local
FP16 graph peak was about 1.84 GB RSS; its quantized counterpart used about
557 MB. ONNX Runtime 1.28 provides an ARM64 wheel with the required
MatMulNBits kernels.
Quantization was compared against the complete FP16 split-graph reference on a
real zoom patch: 49.34 dB PSNR, 0.99986 RGB correlation, and matching edge
energy. A full 40x 2 MP Cloud job passed geometry, color, border, noise, and
detail gates after transferring the model reconstruction residual over the
denoised OEM frame. This preserves real captured edges while still adding VOSR
structure.
Tensor interface
All image and latent tensors use NCHW layout.
| Graph | Inputs | Outputs |
|---|
| DINOv2-L layer 17 | lq_image: FP16 [1,3,512,512] | dino_features: FP16 [1,1024,1024] |
| Qwen VAE encoder | lq_image: FP16 [1,3,512,512]; posterior_noise: FP16 [1,16,64,64] | lq_latent: FP16 [1,16,64,64] |
| DiT prepare | latent_pair: FP16 [1,32,64,64]; timestep: FP32 [1]; next_timestep: FP32 [1]; dino_features: FP16 [1,1024,1024] | hidden: FP16 [1,1024,1536]; conditioning: FP16 [1,1536]; block_conditioning: FP16 [1,9216]; projected_dino: FP16 [1,1024,1536] |
| DiT block stages | hidden, block_conditioning, projected_dino | hidden_out: FP16 [1,1024,1536] |
| DiT final | hidden: FP16 [1,1024,1536]; conditioning: FP16 [1,1536] | velocity: FP16 [1,16,64,64] |
| Qwen VAE decoder | normalized_latent: FP16 [1,16,64,64] | sr_image: FP16 [1,3,512,512] |
Integration outline
Use the Qualcomm AI Engine Direct SDK / QNN runtime to load and execute the DLCs.
The host application is responsible for preprocessing, graph orchestration, random
noise generation, the one-step latent update, tiling, and image postprocessing.
- Resize the low-resolution image to the requested output resolution using bicubic
interpolation, split it into 512 x 512 tiles if needed, convert RGB values to
FP16 NCHW, and normalize them to
[-1, 1].
- Run the DINO graph and Qwen VAE encoder on the same image tile. Supply seeded
normal noise as
posterior_noise if reproducible output is required.
- Create an FP16 normal-noise latent
z with shape [1,16,64,64], concatenate
lq_latent and z along the channel axis, and use the result as latent_pair.
- For the one-step schedule, run DiT prepare with
timestep = [1.0] and
next_timestep = [0.0].
- Pass
hidden sequentially through all four DiT block-stage DLCs, or through
all six sd8g3/qnn-context DiT stages when using the EPContext package.
Reuse block_conditioning and projected_dino for every stage.
- Run DiT final and apply the one-step update
z = z - velocity.
- Decode the updated latent with the Qwen VAE decoder, clamp the output to
[-1, 1], convert it back to RGB, and blend overlapping tiles when tiling.
For 4x super-resolution, bicubic-upscale the source to the final target resolution
before creating the model tiles. The network then restores detail at that target
resolution.
Validation status
- The DINO, VAE encoder, DiT prepare/final, and DiT blocks use their documented
mobile quantization. The VAE decoder deliberately remains FP16.
- All runtime graphs were successfully compiled and linked for SM8650.
- The Qwen VAE encoder and DiT prepare graphs were profiled on the Snapdragon 8
Gen 3 NPU.
- All eleven split EPContext graphs were linked and inspected as DSP v75 / SoC
model 57 contexts. The runtime disables CPU execution-provider fallback.
- The complete on-device chain preserves generated detail: on the S24 Ultra
regression patch, edge energy rose from 1.54 for the bicubic input to 15.19
after VOSR.
- The
00_06 six-block context was profiled successfully on the Galaxy S24
target: 382.9 ms estimated warm inference, 230.9 ms warm load, and 119.5 MB
estimated peak memory. Every reported operator, including the VTCM chunks,
executed on the NPU.
- A monolithic context is intentionally not included. Splitting the DiT into six
VTCM-safe contexts avoids the cloud linker's memory explosion and allows the
application to release each session before opening the next one.
Performance, memory use, and image quality depend on the QNN SDK version, device
firmware, thermal state, tiling implementation, and host-side orchestration. Test on
the exact target device before shipping a production application.
Intended use
This package is intended for research and mobile application development involving:
- photo and screenshot restoration;
- text-rich image enhancement;
- single-image super-resolution;
- Snapdragon NPU deployment experiments.
It is not intended for forensic reconstruction or for recovering information that
is not present in the source image. Generative restoration can introduce plausible
but incorrect details.
Attribution
This is a quantized mobile derivative of VOSR. The original architecture, training,
and checkpoints were created by the VOSR authors. See the
official project and
paper for full details.
Citation
If you use this model, please cite the original VOSR work:
1@inproceedings{wu2026vosr,
2 title = {VOSR: A Vision-Only Generative Model for Image Super-Resolution},
3 author = {Wu, Rongyuan and Sun, Lingchen and Zhang, Zhengqiang and Kong, Xiangtao and Zhao, Jixin and Wang, Shihao and Zhang, Lei},
4 booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
5 year = {2026}
6}
License
Released under the Apache License 2.0, following the upstream VOSR repository.
Users are responsible for reviewing and complying with the licenses and terms of
all upstream components and the Qualcomm QNN SDK/runtime used for deployment.