inference_manifest.json — deployment routing (schema v3)config.json — LeRobot policy config (type=smolvla)model.safetensors — policy torch weights (~865 MB)policy_preprocessor.json + policy_postprocessor.json — normalization stepsHuggingFaceTB/SmolVLM2-500M-Video-Instruct/ — vendored VLM backbone (12 files, ~1.9 GB)artifacts/rknn/rknn_rk3588/ — RKNN compiled modules (5 artifacts)train_config.json — full training hyperparameters| Target | Backend | Runtime | Hardware |
|---|---|---|---|
rknn_rk3588 | rknn | rknn-lite2 | Rockchip RK3588 |
torch-cpu | torch | PyTorch | CPU |
torch-cuda | torch | PyTorch | NVIDIA GPU |
vision_top / vision_wrist (shared vision encoder) -> embedding -> prefill -> action.observation.state [6], observation.current [6], observation.images.top [3,480,640], observation.images.wrist [3,480,640]
Output: action [6] (5 joints + gripper)HuggingFaceTB/SmolVLM2-500M-Video-Instruct/ for offline deployment. The RKNN artifacts were converted from the torch weights. See scripts/train_policy.sh for training and scripts/convert_hmm.sh for conversion procedures.@inproceedings{smolvla,
title = {SmolVLA: Democratizing Cost-Efficient Vision-Language-Action Models for Robot Manipulation},
author = {LeCun, Yann and others},
booktitle = {HuggingFace},
year = {2025}
}
@software{ib_robot,
title = {IB-Robot: Intelligence Boom Robot},
url = {https://gitcode.com/openeuler/IB_Robot},
license = {Apache-2.0}
}