LingBot-Depth for RDK S100/S100P — RGB-D depth refinement
Compiled BPU models (.hbm) for the D-Robotics RDK S100 / S100P, ready to
load — no ONNX export, no calibration, no hb_compile. Built and measured with
BCDL, a C++17 inference and media library
for the RDK S-series with Python bindings.
[!TIP]
Redistributable, including commercially. The licence chain was checked on
the code, the pretrained weights it started from, and the data it was trained
on — all three, because a permissive repository badge does not by itself say
anything about the weights. See Licence.
Files
file
what it is
lingbot_depth_v05_int16_nashm.hbm
all-int16, 480x640 RGB + depth — 1015.3 MB
Measured on an S100P
stage
latency
throughput
refinement
1453 ms
0.69 FPS
hrt_model_exec perf, one thread, minimum of three runs, on a board first gated
against its own previously recorded numbers. BPU time only — CPU
pre/post-processing is on top and is listed per task in BCDL's
benchmark results.
Each task has a decoder in BCDL that turns those raw outputs into boxes,
keypoints, masks, disparity or text — see the
Python API
(中文).
What to know before deploying
This refines depth, it does not estimate it. Give it a depth map you already
have — stereo, ToF, noisy and full of holes — plus the aligned RGB frame, and it
returns hole-filled metric depth with a per-pixel trust mask. It composes with a
stereo or depth camera rather than competing with one.
It is a seconds-per-frame model, an order of magnitude slower than anything
else here. Use it to refine a keyframe on demand — a grasp pose, a mapping
snapshot — not as a stage in a video loop. The cost is the attention score
matrix, quadratic in a sequence of 1 + 2N tokens, moving ~23 GB of DDR per frame.
Only the int16 build is published, and int8 is deliberately absent. int8 PTQ
does not survive a 24-layer ViT-L: it compiles cleanly and returns a well-formed
depth map whose dynamic range has collapsed — 2.9-13.8 m against the float
model's 0.97-45.8 m, 233% mean absolute relative error. Publishing it would just
be handing somebody a trap.
The deployed graph also keeps every depth token, where upstream drops the ones
whose patch holds no valid reading — that masking makes the sequence length
depend on the data and cannot be compiled statically. Measured cost of keeping
them: 0.06% mean absolute relative error, 0.9999 mask IoU, on scenes that are
87-100% valid. Very sparse input depth was not measured.
Licence
Apache-2.0 on the code and the weights; the DINOv2 backbone is Apache-2.0 too.
BCDL itself is Apache-2.0 and is unrelated to these terms — it is a
general-purpose runtime that loads any .hbm. The licence above constrains
these weights and this compiled artefact.
The conversion recipe — ONNX export, calibration, hb_compile config and the
acceptance numbers — is public in
bcdl-model-zoo, so this build can
be reproduced or retargeted rather than taken on trust.