Diffusion Policy — banana-in-pot (JOINT, fp16)
Diffusion Policy trained on the "put the right banana in the pot" task (UR7e + GELLO
teleoperation, 2 RGB cameras), in JOINT action space (6 joints + gripper), using
fp16 mixed-precision training.
- Checkpoint: step 80,000 (best open-loop MAE)
- Base library: LeRobot 0.6.1 (pin
8a74e0a)
- Dataset:
Bigenlight/banana_in_pot_lerobot_v3 — 51 episodes / 21,524 frames / 30 fps
- fp32 sibling:
Bigenlight/diffusion_banana_in_pot_joint
Architecture
ResNet18 (ImageNet-pretrained, trained end-to-end) + SpatialSoftmax vision encoder →
FiLM-conditioned 1D temporal U-Net denoiser (DDPM-100, epsilon prediction). ~278M params.
Images resized to 360×640, crop off. n_obs_steps=2, horizon=64, n_action_steps=32.
Training
- Precision: fp16 via HF Accelerate
mixed_precision=fp16 (automatic GradScaler).
- Requires a
dtype field on DiffusionConfig (absent upstream at this pin); launched with
--policy.dtype=float16.
- Batch 8, 80k steps, AdamW, seed 1000, 45 train / 6 held-out episodes.
- Hardware: single RTX A4000. ~4.48 step/s, ~4.96 h (vs ~3.49 step/s / ~6.4 h fp32 →
~1.25× faster, ~22% less wall-clock). No NaN/instability.
Open-loop evaluation (DDIM-10, held-out episodes 45–50)
| fp32 @80k | fp16 @80k |
|---|
| poseMAE (rad) | 0.0845 | 0.08268 |
| gripper accuracy | 0.953 | 0.960 |
fp16 matches (marginally beats, within noise) fp32 accuracy at ~22% lower training cost.
Select the deploy checkpoint by open-loop MAE, not eval_loss (which rises during
training for diffusion policies without indicating overfitting).
Intended use & limitations
Research artifact. Small single-task, single-scene, real-world (noisy) dataset of 51
success-only demonstrations; offline metrics only — no closed-loop hardware success rate
measured yet. Not safety-validated for autonomous operation.