Views
No views yet
LangForce has been accepted to ICML 2026, and you can find our ckpt in huggingface.LangForce has been integrated into starVLA. You can now directly train LangForce through starVLA and perform end-to-end training and evaluation on benchmarks such as LIBERO, SimplerEnv, and RoboCasa.Information Collapse. Consequently, models degenerate into vision-only policies that ignore language constraints and fail in out-of-distribution (OOD) settings. To address this, we propose LangForce:, a novel framework that enforces instruction following via Bayesian decomposition. By introducing learnable Latent Action Queries, we construct a dual-branch architecture to estimate both a vision-only prior $p(a \mid v)$ and a language-conditioned posterior $\pi(a \mid v, \ell)$. We then optimize the policy to maximize the conditional Pointwise Mutual Information (PMI) between actions and instructions. This objective effectively penalizes the vision shortcut and rewards actions that explicitly explain the language command. Without requiring new data, LangForce significantly improves generalization. Extensive experiments across on SimplerEnv and RoboCasa demonstrate substantial gains, including an 11.3% improvement on the challenging OOD SimplerEnv benchmark, validating the ability of our approach to robustly ground language in action.| Method | SimplerEnv (Avg) | RoboCasa (Avg) | LIBERO (Avg) | VLA-Arena (L0 / L1 / L2 Avg) |
|---|---|---|---|---|
| QwenGR00T (Baseline) | 55.2% | 47.8% | 96.5% | 76.9% / 23.5% / 12.5% |
| LangForce (Ours) | 66.5% (+11.3%) | 52.6% (+4.8%) | 98.4% (+1.9%) | 84.2% (+7.3%) / 38.8% (+15.4%) / 24.1% (+11.5%) |
Values in each model cell are ordered as L0 / L1 / L2. Bold denotes the best result and italics denotes the second-best result within each metric and difficulty level. Higher is better for SR; lower is better for CC. Ties receive the same formatting.~~0.00~~highlights task failures.
| Dimension | Task | Metric | π0.5 | GR00T-N1.6 | Qwen3-VL-OFT | Qwen3-VL-GR00T | Qwen3-VL-PI | LingBot-VLA | Motus | LangForce |
|---|---|---|---|---|---|---|---|---|---|---|
| Safety | Static Obstacles | SR | 0.90 / 0.62 / 0.40 | 0.72 / 0.30 / 0.14 | 0.84 / 0.14 / 0.08 | 0.91 / 0.18 / 0.05 | 0.62 / 0.24 / 0.16 | 0.47 / | 0.82 / 0.56 / 0.19 | 0.95 / 0.72 / 0.56 |
| Safety | Static Obstacles | CC | 0.00 / 33.3 / 76.6 | 0.00 / 11.2 / 38.4 | 0.00 / 16.1 / 11.8 | 0.00 / 9.1 / 29.6 | 0.00 / 14.3 / 21.4 | 25.2 / | 0.00 / 13.3 / 55.7 | 0.00 / 18.9 / 34.9 |
| Safety | Cautious Grasp | SR | 0.50 / 0.14 / | 0.16 / 0.02 / | 0.86 / 0.04 / | 0.77 / 0.08 / 0.02 | 0.80 / 0.20 / | 0.21 / | 0.76 / 0.07 / 0.19 | 0.75 / 0.20 / 0.07 |
| Safety | Cautious Grasp | CC | 5.0 / 5.5 / 1.2 | 9.6 / 41.2 / 10.4 | 7.1 / 40.3 / 6.3 | 2.4 / 98.2 / 11.1 | 2.9 / 99.3 / 19.6 | 10.5 / | 3.7 / 62.2 / 15.3 | 1.7 / 24.8 / 16.6 |
| Safety | Hazard Avoidance | SR | 0.58 / 0.30 / 0.36 | 0.64 / 0.04 / 0.10 | 0.50 / 0.12 / 0.04 | 0.67 / 0.15 / 0.19 | 0.54 / 0.16 / 0.14 | 0.16 / | 0.43 / 0.24 / 0.16 | 0.74 / 0.28 / 0.40 |
| Safety | Hazard Avoidance | CC | 7.1 / 15.0 / 14.5 | 6.1 / 20.2 / 17.9 | 8.9 / 20.9 / 22.2 | 7.0 / 19.3 / 19.5 | 9.6 / 19.3 / 20.0 | 25.2 / | 6.8 / 23.6 / 27.8 | 0.60 / 5.0 / 3.8 |
| Safety | State Preservation | SR | 0.58 / 0.56 / 0.54 | 0.66 / 0.50 / 0.38 | 0.90 / 0.40 / 0.30 | 0.86 / 0.56 / 0.47 | 0.78 / 0.50 / 0.52 | 0.54 / | 0.85 / 0.47 / 0.49 | 0.90 / 0.66 / 0.31 |
| Safety | State Preservation | CC | 0.00 / 5.6 / 20.8 | 0.00 / 5.0 / 10.4 | 0.00 / 4.0 / 5.2 | 0.00 / 5.6 / 15.7 | 0.00 / 5.0 / 9.6 | 5.4 / 0.00 / | 0.00 / 4.7 / 18.9 | 0.00 / 3.6 / 4.4 |
| Safety | Dynamic Obstacles | SR | 0.50 / 0.44 / 0.22 | 0.74 / 0.50 / 0.02 | 0.64 / 0.48 / 0.08 | 0.81 / 0.56 / 0.03 | 0.60 / 0.36 / 0.12 | 0.40 / | 0.43 / 0.35 / 0.26 | 0.91 / 0.68 / 0.36 |
| Safety | Dynamic Obstacles | CC | 2.4 / 8.8 / 5.7 | 5.7 / 7.3 / 56.8 | 5.6 / 12.2 / 7.6 | 6.0 / 8.3 / 2.7 | 4.1 / 2.9 / 11.2 | 27.8 / | 3.5 / 39.6 / 35.3 | 1.1 / 28.7 / 36.9 |
| Distractor | Static Distractors | SR | 0.88 / 0.16 / 0.16 | 0.46 / 0.32 / 0.06 | 0.82 / 0.06 / 0.02 | 0.91 / 0.06 / 0.02 | 0.80 / 0.14 / 0.02 | 0.93 / 0.15 / 0.11 | 0.75 / 0.19 / 0.03 | 0.93 / 0.26 / 0.04 |
| Distractor | Dynamic Distractors | SR | 0.80 / 0.66 / 0.54 | 0.70 / 0.72 / 0.18 | 0.82 / 0.48 / 0.20 | 0.92 / 0.49 / 0.25 | 0.90 / 0.56 / 0.30 | 0.88 / 0.61 / 0.17 | 0.73 / 0.60 / 0.33 | 0.91 / 0.72 / 0.45 |
| Extrapolation | Preposition Combinations | SR | 0.62 / 0.24 / 0.06 | 0.48 / | 0.54 / | 0.51 / 0.01 / | 0.38 / | 0.46 / 0.05 / 0.01 | 0.13 / | 0.75 / 0.03 / 0.02 |
| Extrapolation | Task Workflows | SR | 0.38 / 0.20 / 0.22 | 0.42 / | 0.42 / 0.20 / 0.16 | 0.51 / 0.03 / 0.09 | 0.46 / 0.02 / 0.10 | 0.37 / 0.05 / 0.11 | 0.32 / | 0.63 / 0.11 / 0.19 |
| Extrapolation | Unseen Objects | SR | 0.48 / 0.60 / 0.20 | 0.26 / 0.18 / 0.16 | 0.60 / 0.24 / 0.06 | 0.63 / 0.46 / 0.26 | 0.40 / 0.60 / 0.06 | 0.34 / 0.32 / 0.15 | 0.59 / 0.55 / 0.13 | 0.80 / 0.61 / 0.25 |
| LongHorizon | Long Horizon | SR | 0.85 / | 0.29 / 0.02 / | 0.98 / | 0.96 / | 0.76 / | 0.82 / 0.03 / | 0.64 / 0.05 / 0.03 | 0.99 / |
1# Clone the repo
2git clone https://github.com/starVLA/starVLA
3
4# Create conda environment
5conda create -n starVLA python=3.10 -y
6conda activate starVLA
7
8# Install requirements
9pip install -r requirements.txt
10
11# Install FlashAttention2
12pip install flash-attn --no-build-isolation
13
14# Install starVLA
15pip install -e .torch==2.6.0+cu12.4
flash-attention==2.7.4.post1
## If using Qwen3.5 as the VLM
flash-linear-attention==0.3.2
causal_conv1d==1.5.0.post8add_token.py to update the tokenizer with these additional tokens.1conda activate starvla
2cd /xxx/worlkplace/starVLA-v2.0
3
4export NCCL_SOCKET_IFNAME=eth0
5export NCCL_IB_DISABLE=1
6export NCCL_BLOCKING_WAIT=1
7export NCCL_ASYNC_ERROR_HANDLING=1
8export NCCL_TIMEOUT=1000 # timeout set to 1 hour (unit: seconds)
9
10framework_name=LangForceV5
11base_vlm=/xxx/starVLA-v2.0/playground/Pretrained_models/Qwen3-VL-4B-with-Action-Query
12run_id=GR00T_Simpler_LangForce
13freeze_module_list=''
14config_yaml=./examples/SimplerEnv/train_files/starvla_cotrain_oxe.yaml
15oxe_data_root=/xxx/starVLA/playground/Datasets/OXE_LEROBOT_DATASET/
16data_mix=bridge
17run_root_dir=./results/LangForce/SimplerEnv
18
19output_dir=${run_root_dir}/${run_id}
20mkdir -p ${output_dir}
21
22accelerate launch \
23 --config_file starVLA/config/deepseeds/deepspeed_zero2.yaml \
24 --num_processes 8 \
25 starVLA/training/train_starvla.py \
26 --config_yaml ${config_yaml} \
27 --framework.name ${framework_name} \
28 --framework.qwenvl.base_vlm ${base_vlm} \
29 --framework.qwenvl.template ${vlm_template} \
30 --framework.detach_prior_cond ${detach_prior_cond} \
31 --framework.qwenvl.num_latent_action_query ${num_latent_action_query} \
32 --framework.action_model.diffusion_model_cfg.num_layers ${dit_num_layers} \
33 --datasets.vla_data.data_root_dir ${oxe_data_root}\
34 --datasets.vla_data.data_mix ${data_mix} \
35 --datasets.vla_data.per_device_batch_size ${per_device_batch_size} \
36 --trainer.freeze_modules ${freeze_module_list} \
37 --trainer.max_train_steps 100000 \
38 --trainer.save_interval 10000 \
39 --trainer.logging_frequency 100 \
40 --trainer.eval_interval 1000 \
41 --run_root_dir ${run_root_dir} \
42 --run_id ${run_id} \
43 --wandb_project starVLA \
44 --wandb_entity xxxLangForce is currently under active development. Feel free to check back frequently for updates and new features!
1@inproceedings{LangForce_2026_ICML,
2 title = {LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries},
3 author = {Lian, Shijie and Yu, Bin and Lin, Xiaopeng and Yang, Laurence T. and Shen, Zhaolong and Wu, Changti and Miao, Yuzhuo and Huang, Cong and Chen, Kai},
4 booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
5 year = {2026},
6 series = {Proceedings of Machine Learning Research},
7 publisher = {PMLR},
8 url = {https://arxiv.org/abs/2601.15197}
9 }