This dataset is associated with the paper "Recognizing Co-speech Gestures in-the-Wild" (ECCV 2026).
Our aim is to recognise and localize semantic gestures in real-world videos. These gestures are visually depictive and semantically linked to specific spoken words. We introduce a new large-scale benchmark, GRW (Gesture Recognition in-the-Wild), which provides word-level annotations and gesture boundaries for semantic… See the full description on the dataset page:
https://huggingface.co/datasets/sindhuhegde/grw.