Dataset generated by merging the SignTalk-GH corpus with the GSL dictionary (zenodo openpose videos) and sampling to a balanced subset. Intended for downstream preprocessing (preprocessing.ipynb) and Text2Sign pipelines.
Source corpora: SignTalk-GH (Videos + Metadata.xlsx) and GSL OpenPose data.
Integration: GSL concepts normalized (token splitting on
OR,
AND, hyphens preserved)… See the full description on the dataset page:
https://huggingface.co/datasets/zahemen9900/signtalk-gh-sampled.