This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Video ID and link… See the full description on the dataset page:
https://huggingface.co/datasets/dm-petrov/youtube-commons-small.