Views
No views yet
nano-vla-arm,
but the scene is rendered through a fixed perspective homography — a receding
tabletop, foreshortened blocks — so the policy must read a camera-like image, not a
flat schematic. Run it live at https://physicalai-bmi.org/research/vla.| metric | value |
|---|---|
| Reaches correct block, trained instructions | 97.0% |
| Reaches correct block, novel instructions | 97.0% |
| Correct when instruction flipped | 0.0% |
config.Hinv) is applied
identically in training and in the browser, so train == live. A small reference
policy on a rendered task; not a foundation model.model.safetensors, vla.web.json (float32 for in-browser, verified 2.3e-7
vs safetensors), metrics.json. CC-BY-4.0, Institute for Physical AI @ BMI.