Views
No views yet
CompiledModel API (ML Drift delegate). No CPU fallback ops — the whole graph is GPU-compatible.| Task | Monocular depth (single RGB → depth) |
| Backbone | DINOv2 ViT-S + RoPE, DPT/DualDPT depth head |
| Input | [1, 3, 896, 504] NCHW float32, ImageNet-normalized, native portrait aspect |
| Output | [1, 1, 896, 504] depth |
| Precision / size | FP16, 55 MB |
| Device | Pixel 8a, LiteRT GPU (Accelerator.GPU), ~1.8 s / image |
| Fidelity | Pearson corr 0.99948 vs the official PyTorch DA3-Small pipeline |
upper_bound_resize, longer side → 896, multiple of 14).
Forcing a square 896×896 and letterbox-padding drops the match to corr 0.977 (the black padding leaks into
the content through global attention). Converting at the native rectangle restores corr 0.9994 and is also
faster (fewer tokens). This checkpoint is built for portrait ~9:16. For another aspect, re-convert at that
shape (or your camera's fixed aspect) with the script below.resize to 504×896 (W×H) → x/255 → (x - mean) / std
mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225] # ImageNet, RGB, NCHWlitert-torch. DA3 is not GPU-clean out of the box; the following exact, GPU-clean rewrites
were applied (all numerically faithful unless noted):model. key-prefix strip (load fix)max_position = int(positions.max())+1 → constant (torch.export data-dependent)gamma folded into attn.proj / mlp.fc2 (the LayerScale MUL otherwise mis-lays-out the
token dim on the GPU delegate: fully_connected {1,1,N,C} vs {N,1,1,C})pos_embed bicubic interpolation baked to a constant (the interpolate of a constant emits GATHER_ND
on desktop and RESIZE_BILINEAR with 0 runtime inputs on device)Conv2d (flipped
weight) — exact equivalent (~1e-7), because the Pixel-8a GPU rejects TRANSPOSE_CONV and the conv+
depth-to-space alternative needs >4Dcustom_interpolate align_corners=True → False (GPU bans align_corners=True resize) — the
only non-exact rewrite; source of the residual ~0.05 % vs the official modelmake_sincos broadcast emits BROADCAST_TO; ratio-0.1 refinement)x[:, :, 0] = cam_token → torch.cat (in-place index-assign → SELECT_V2)GATHER_ND = 0, no >4D tensors, no TRANSPOSE_CONV / BROADCAST_TO / banned ops.align_corners=True→False change in (7), which the mobile GPU forces — an irreducible
hardware constraint, not a conversion error. Structure and edge sharpness are visually identical.1val model = CompiledModel.create(context.assets, "da3_small_gpu_fp16.tflite",
2 CompiledModel.Options(Accelerator.GPU), null)
3// input: [1,3,896,504] NCHW, ImageNet-normalized; output: [1,1,896,504] depth