Views
No views yet
| Category | Details |
|---|---|
| Train | 8 × H100 GPUs, each with 80GB VRAM (batch size: 256) |
| Model size | 1.3B (CLIP-RT base + 0.3B action decoder) |
| Action dimension | 7D end-effector action × 8 action chunks |
| Loss | L1 regression |
| Epochs | 128 |
| Performance | 95.2% success rate on the LIBERO-Spatial task suite |
| Throughput | 163Hz |
| Inference | One GPU with 9GB VRAM |
1@article{kang2024cliprt,
2 title={CLIP-RT: Learning Language-Conditioned Robotic Policies from Natural Language Supervision},
3 author={Kang, Gi-Cheon and Kim, Junghyun and Shim, Kyuhwan and Lee, Jun Ki and Zhang, Byoung-Tak},
4 journal={arXiv preprint arXiv:2411.00508},
5 year = {2024}
6}