This is our pick and place model trained for the
ROS2SmolVLA project.
It has been trained on this dataset:
ROS2SmolVLA_ur10e_no_joints_crop_pick_place
The top camera image has been cropped to 720x720 and the observations only contain the cartesian pose of the endeffector.
"linear_x.vel",
"linear_y.vel",
"linear_z.vel",
"angular_x.vel",
"angular_y.vel",
"angular_z.vel",
"gripper.pos"