StarVLA QwenPI_v3 Qwen3-VL-4B for RoboDojo (100k)
This directory contains one StarVLA QwenPI_v3 checkpoint initialized from
Qwen3-VL-4B-Instruct and trained on the 35-task RoboDojo LeRobot v2.1 mixture.
The VLM, VLM interface, and action model were trained end to end; this is not a
LoRA or adapter-only checkpoint.
Model details
| Item | Value |
|---|
| Framework | StarVLA QwenPI_v3 |
| Base VLM | Qwen3-VL-4B-Instruct |
| Action model | 36-layer LayerwiseFM |
| Action representation | 14D absolute joint position (abs_qpos) |
| Action horizon | 50 |
| State dimension | 14 |
| Camera input | Head, left wrist, right wrist; resized to 224 x 224 |
| Inference flow steps | 4 |
| Checkpoint step | 100,000 |
| Checkpoint format | Complete StarVLA framework state dict (.pt) |
The policy predicts a normalized 50 x 14 action chunk. RoboDojo evaluation
must use the saved arx_x5 normalization statistics and execute 16 actions
before requesting the next chunk.
Files
1README.md
2config.yaml
3config.full.yaml
4dataset_statistics.json
5summary.jsonl
6checkpoints/
7└── steps_100000_pytorch_model.pt
Only the requested 100k checkpoint is included. Keep the configuration and
dataset statistics beside the checkpoints/ directory; StarVLA uses them to
reconstruct the framework and unnormalize actions.
Training details
| Setting | Value |
|---|
| Dataset mixture | robodojo_v21_all_h50_q99 |
| Training tasks | 35 |
| Per-GPU batch size | 16 |
| Gradient accumulation | 1 |
| Frozen modules | None |
| Optimizer | AdamW, betas (0.9, 0.95), epsilon 1e-8 |
| VLM learning rate | 1e-5 |
| VLM-interface learning rate | 1e-5 |
| Action-model learning rate | 1e-4 |
| Schedule | Cosine, 5,000 warmup steps, minimum LR 5e-7 |
| Gradient checkpointing | Enabled |
| Random seed | 42 |
Official RoboDojo evaluation
All policies below use the official complete 42-task protocol: 50 episodes per
task, 2,100 episodes per policy. Values are shown as SR (%) / Score. This
directory's policy is bolded. Higher is better for both SR and Score.
Group summary
| Policy | Average | Generalization | Precision | Long-Horizon | Memory | Open |
|---|
| QwenOFT | 4.86 / 8.01 | 4.33 / 6.42 | 11.75 / 17.54 | 5.50 / 12.95 | 1.67 / 1.77 | 0.50 / 0.60 |
| QwenGR00T | 3.81 / 7.35 | 3.50 / 6.52 | 5.75 / 10.09 | 6.50 / 15.46 | 3.33 / 4.37 | 0.00 / 0.00 |
| QwenPI_v3 | 6.19 / 9.60 | 4.17 / 7.28 | 14.00 / 19.06 | 10.00 / 17.84 | 2.00 / 2.32 | 0.75 / 0.88 |
Task details
Task values are SR (%) / Score; each task uses 50 episodes. The bolded
column is the policy in this directory.
| Evaluation group / task | QwenOFT | QwenGR00T | QwenPI_v3 |
|---|
| Generalization | 4.33 / 6.42 | 3.50 / 6.52 | 4.17 / 7.28 |
| stack_bowls | 18.00 / 21.00 | 10.00 / 14.80 | 14.00 / 16.70 |
| push_T | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| pack_objects_into_box | 0.00 / 3.10 | 0.00 / 7.80 | 0.00 / 6.80 |
| fold_clothes | 10.00 / 12.80 | 8.00 / 12.40 | 2.00 / 9.60 |
| hang_mugs | 0.00 / 3.60 | 0.00 / 3.00 | 0.00 / 3.50 |
| sweep_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 2.00 / 2.00 |
| pour_liquid_into_cup | 14.00 / 14.00 | 14.00 / 14.00 | 12.00 / 12.00 |
| make_toast | 0.00 / 1.00 | 2.00 / 5.00 | 2.00 / 5.00 |
| arrange_largest_number | 0.00 / 1.90 | 2.00 / 4.10 | 2.00 / 5.70 |
| sort_nesting_dolls_by_size | 0.00 / 0.00 | 4.00 / 4.00 | 6.00 / 6.00 |
| store_laptop_and_headphones | 4.00 / 11.20 | 0.00 / 7.20 | 2.00 / 8.40 |
| stack_blocks | 6.00 / 8.40 | 2.00 / 5.90 | 8.00 / 11.60 |
| Precision | 11.75 / 17.54 | 5.75 / 10.09 | 14.00 / 19.06 |
| fasten_screws | 4.00 / 8.00 | 0.00 / 2.00 | 0.00 / 6.00 |
| plug_in_charger | 6.00 / 6.00 | 2.00 / 2.00 | 4.00 / 4.00 |
| insert_tubes | 40.00 / 51.60 | 28.00 / 40.40 | 44.00 / 56.80 |
| pour_balls_into_vase | 8.00 / 8.00 | 0.00 / 0.00 | 2.00 / 2.00 |
| play_Xylophone | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| deposit_coin | 0.00 / 3.20 | 2.00 / 5.60 | 6.00 / 7.60 |
| insert_key | 0.00 / 12.90 | 0.00 / 9.90 | 0.00 / 11.10 |
| build_tower | 36.00 / 50.60 | 14.00 / 20.80 | 56.00 / 65.00 |
| Long-Horizon | 5.50 / 12.95 | 6.50 / 15.46 | 10.00 / 17.84 |
| put_bottles_into_dustbin | 22.00 / 40.90 | 26.00 / 44.40 | 64.00 / 73.60 |
| fill_pen_holder | 4.00 / 11.70 | 6.00 / 14.40 | 8.00 / 23.00 |
| classify_objects | 2.00 / 5.50 | 6.00 / 11.50 | 0.00 / 7.50 |
| play_tic_tac_toe | 0.00 / 12.40 | 2.00 / 16.40 | 0.00 / 6.80 |
| fill_egg_holder | 0.00 / 0.60 | 0.00 / 0.00 | 0.00 / 0.80 |
| organize_table | 0.00 / 16.50 | 0.00 / 25.00 | 4.00 / 27.00 |
| make_kong | 16.00 / 16.00 | 12.00 / 12.00 | 4.00 / 4.00 |
| play_stacking_toy | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| Memory | 1.67 / 1.77 | 3.33 / 4.37 | 2.00 / 2.32 |
| cover_blocks | 0.00 / 0.60 | 6.00 / 12.10 | 0.00 / 1.50 |
| match_and_pick_from_conveyor | 10.00 / 10.00 | 14.00 / 14.00 | 12.00 / 12.00 |
| swap_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| swap_T | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| press_by_number | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| imitate_sorting_sequence | 0.00 / 0.00 | 0.00 / 0.10 | 0.00 / 0.40 |
| Open | 0.50 / 0.60 | 0.00 / 0.00 | 0.75 / 0.88 |
| align_blocks | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| general_pickup | 4.00 / 4.00 | 0.00 / 0.00 | 6.00 / 6.00 |
| stack_blocks_by_language | 0.00 / 0.80 | 0.00 / 0.00 | 0.00 / 0.80 |
| solve_equation | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| classify_objects_by_language | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.20 |
| pick_from_conveyor_by_image | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| store_tools_in_toolbox | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| pour_by_language | 0.00 / 0.00 | 0.00 / 0.00 | 0.00 / 0.00 |
Evaluation
This is a StarVLA checkpoint, not a Hugging Face from_pretrained() directory.
Use the StarVLA model server and the XPolicyLab StarVLA adapter.
Start the model server from the StarVLA repository:
1huggingface-cli download StarVLA/StarVLA-Qwen3vl4b-PIv3-RoboDojo \
2 --local-dir StarVLA-Qwen3vl4b-PIv3-RoboDojo
3export CKPT="$PWD/StarVLA-Qwen3vl4b-PIv3-RoboDojo/checkpoints/steps_100000_pytorch_model.pt"
4python deployment/model_server/server_policy.py \
5 --ckpt_path "$CKPT" \
6 --port 57700 \
7 --use_bf16
Run a RoboDojo task from XPolicyLab/policy/starVLA:
1STARVLA_CKPT_PATH="$CKPT" \
2STARVLA_INCLUDE_STATE=True \
3STARVLA_UNNORM_KEY=arx_x5 \
4STARVLA_EXECUTE_HORIZON=16 \
5bash eval.sh \
6 RoboDojo build_tower qwenpi_v3_steps_100000 \
7 arx_x5 joint 0 0 1 <policy_conda_env> <robodojo_conda_env>
The final arguments are the seed, policy GPU, simulator GPU, policy environment,
and RoboDojo environment. Use the RoboDojo task registry's native episode counts
when producing an official aggregate.
Evidence and evaluation boundary
- Architecture, input/action dimensions, normalization contract, training
settings, and the 100k release were checked against
config.full.yaml,
dataset_statistics.json, and the Hub file tree.
- The official-protocol table is retained from this published reference Card
and matches the QwenPI_v3 100k checkpoint. This repository does not include
raw per-episode evaluation logs, so the aggregate is not independently
reconstructable from the repository files alone.
- Rows for QwenOFT and QwenGR00T are comparison context; they are not results
produced by the checkpoint in this repository.
- Evaluation requires the same
arx_x5 statistics, state inclusion, camera
order, absolute-joint action mapping, and 16-step execution horizon.
Intended use and limitations
This checkpoint is intended for RoboDojo simulation research with the ARX X5
dual-arm embodiment. Performance with different camera calibration, state/action
ordering, normalization, robots, or real-world hardware has not been established.
License evidence
The Apache-2.0 metadata above is retained from this target repository's
previously published Model Card; it is not inferred from the StarVLA code
license. No separate LICENSE file is packaged, and base-model and dataset
terms remain applicable.