Views
No views yet
| Paper row | Table 17 / Figure 4: RoboTwin 2.0 'lift pot' Drop-12 |
| Dropped blocks | Language backbone (PaliGemma, 18 layers): drop 12 whole blocks selected by per-task GateProbe; keep blocks [0,1,2,3,4,5]. Vision and action untouched. |
| Recovery training | single-task RoboTwin 2.0, FLOPs-matched to baseline (step 34234), lr 5e-5 |
| RoboTwin 2.0 success rate | Easy 19 / Hard 4 |
1python scripts/serve_policy_batch_drop.py \
2 --config pi05_libero_dropped \
3 --dir <this_repo_local_path> \
4 --port 8000llm_drop_attn_list / llm_drop_mlp_list shown above (via config or CLI) when serving,
otherwise layers will be mismatched. assets/ contains the LIBERO norm stats.
The optimizer state (train_state/) is not included.1@article{sun2026vladrop,
2 title={Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?},
3 author={Sun, Guoheng and Feng, Kaixi and He, Shwai and Gong, Xiaochuan and He, Yexiao and Wang, Ziyao and Shen, Zheyu and Ye, Wanghao and Kompella, Ramana Rao and Liu, Gaowen and Li, Ang},
4 journal={arXiv preprint arXiv:2606.27755},
5 year={2026}
6}