Views
No views yet
| Field | Value |
|---|---|
| Architecture | ACT (ResNet18 backbone + 4-layer Transformer encoder + VAE chunking head) |
| Dataset | JHeisler/aloha_solo_left_4_6_26 — 50 episodes, 29,785 samples, 30 fps |
| State / action dim | 9 / 9 |
| Cameras | cam_high, cam_left_wrist (3×480×640 each) |
| Steps | 40,000 |
| Batch size | 48 |
| Learning rate | 6e-5 (linear warmup 500 → cosine) |
| Total samples seen | ~1.92M (~64 epochs over the dataset) |
| AMP | enabled |
| torch.compile | enabled |
| Save freq | every 10,000 steps (10k / 20k / 30k / 40k checkpoints) |
| Final loss | ~0.015 |
| Final grad norm | ~0.19 |
| Wall clock | ~6h 7min on RTX A4500 |
| LeRobot pin | 96c7052777aca85d4e55dfba8f81586103ba8f61 |
| Workstream | Model | Steps | Samples | HF |
|---|---|---|---|---|
| S001 | ACT | 13,400 | 640K | act_left |
| S002 | Hybrid ACT+Diffusion | 13,400 | 321K | act_diffusion |
| S003 | ACT (shipped) | 40,000 | 1.92M | this repo |
| S004 | Hybrid ACT+Diffusion | 40,000 | 1.12M | act_diffusion_40k |
1from lerobot.common.policies.act.modeling_act import ACTPolicy
2policy = ACTPolicy.from_pretrained("JHeisler/aloha_solo_left_4_6_26_act_left_40k")96c7052.