Real-robot trajectories for "put the toast in the plate" with human pairwise preference
labels on multiple judgment axes. Built for reward-model / preference-learning
research: every label is a comparison of two trajectories on one named axis, not a
scalar score.
The trajectory data is a standard LeRobot
v2.1 dataset, so it also loads directly as an imitation-learning dataset.