openpi pi05_bi checkpoint for a bimanual two-tube pick-and-place task with
tactile inputs. Step 18000 (3.19 epochs), held-out validation
loss 0.0563.
This checkpoint is past the optimum and is not recommended for use. The best of the run is step 14000 at 0.0537 (Aether258/pi05_bi_two_tubes_0102_step14000). This one is published as a reference point for what the model looks like after generalization has degraded.
Task
Single unified instruction for every episode:
Firstly, use the left hand to pick up the blue tube, and then use the right
hand to pick up the green tube. Next, use the left hand to place the blue tube
back firstly, and then use the right hand to place the green tube back.
Data
source
episodes
frames
KaiyueChen/two_tubes_01
519
425,115
KaiyueChen/two_tubes_02
500
377,604
merged
1,019
802,719
LeRobot v2.1, 30 fps, robot_type=bimanual, images embedded in the parquet
files. Six camera streams: camera0, camera1, and four tactile sensors
(tactile_left_0/1, tactile_right_0/1).
Split, held out per source repo (10%, seed 42): train 917 / val_seen 102 /
val_unseen 102. Normalization statistics (quantile q01/q99) over train only.
(CosineDecaySchedule defaults -- pi05_bi does not override lr_schedule)
LoRA
rank 16 on the LLM, rank 32 on the action expert
vision tower
fully fine-tuned -- the freeze filter matches only .*llm.*
The run spans three hosts, chained with --resume at steps 10000 and 16000.
--resume restores weights and optimizer state but not the data loader
position, so the batch sequence after each restart differs from what an
uninterrupted run would have seen.
Validation curve
Flow-matching loss, 20 batches per split, same leading batches every evaluation.
step
train
val_seen
val_unseen
gap
run
0
0.5525
0.4968
0.5261
0.0293
hvxln3lg
2000
0.0553
0.0504
0.0608
0.0104
hvxln3lg
4000
0.0490
0.0467
0.0576
0.0109
hvxln3lg
6000
0.0460
0.0437
0.0543
0.0106
hvxln3lg
8000
0.0441
0.0423
0.0550
0.0127
hvxln3lg
10000
0.0435
0.0416
0.0542
0.0126
hvxln3lg / resume point
12000
0.0420
0.0403
0.0538
0.0135
q9ymwova
14000
0.0404
0.0387
0.0537
0.0150
q9ymwova
16000
0.0395
0.0383
0.0551
0.0168
q9ymwova / resume point
18000
0.0383
0.0374
0.0563
0.0189
i5s1jqvg
20000
0.0376
0.0363
0.0567
0.0204
i5s1jqvg
22000
0.0361
0.0363
0.0584
0.0221
i5s1jqvg
val_unseen bottomed at step 14000 (0.0537) and rose monotonically afterwards --
0.0551, 0.0563, 0.0567, 0.0584 -- while val_seen kept improving and the gap grew
monotonically from 0.0150 to 0.0220. Five consecutive points moving the same
direction; this is unambiguous overfitting, not the noisy plateau seen between
steps 6000 and 14000. Training was stopped at step ~23400.
The step-16000 row is taken from the training host's local log: that host was shut
down before wandb flushed its final metrics, so the cloud run q9ymwova ends at
step 15800.
Note the training loss is measured on augmented images (random crop to 95%,
+-5 deg rotation, colour jitter, on all six streams including the four tactile
ones) while validation runs with train=False, which skips augmentation. The two
columns are therefore not directly comparable.
Contents
checkpoint/
params/ # inference weights
train_state/ # optimizer state, for resuming
assets/two_tubes_0102/
norm_stats.json # computed over the train split only