SmolVLA Fine-tuned on LIBERO-Plus (Object)
This model is a fine-tuned version of lerobot/smolvla_libero specifically trained on the LIBERO-Plus task suite (libero_object). It is designed to perform robotic manipulation tasks by taking visual observations and language instructions as input to predict 7-DoF actions.
🤖 Model Details
- Model Type: Vision-Language-Action (VLA)
- Base Model: lerobot/smolvla_libero
- Task Suite: LIBERO-Plus (libero_object)
- Framework: LeRobot
- Action Space: 7-DoF (dx, dy, dz, droll, dpitch, dyaw, gripper)
- Language: English
📊 Training Information
The model was trained by freezing the vision and language encoders and fine-tuning only the action expert modules and state projection layers.
- Dataset:
lerobot/libero_plus (libero_object)
- Training Steps: 10,000 steps
- Batch Size: 8
- Learning Rate: 1e-4 (Cosine decay with warmup)
- Hardware: Single NVIDIA L4 GPU
🚀 How to Use
You can load and use this model directly with the Hugging Face lerobot library.
Installation