BARISTA is a densely annotated egocentric video dataset of coffee preparation, designed for unified benchmarking of vision-language models across spatial, temporal, relational, and procedural understanding tasks.
The dataset contains 185 egocentric videos (~4.4 hours, 30 FPS, 1280×720 to 1920×1080) covering three coffee preparation methods: capsule machines, portafilter machines, and fully automatic machines. Videos were recorded in controlled indoor setups using iPhones… See the full description on the dataset page:
https://huggingface.co/datasets/ramblr/BARISTA.