Anim-400K: A dataset designed from the ground up for automated dubbing of video
What is Anim-400K?
Anim-400K is a large-scale dataset of aligned audio-video clips in both the English and Japanese languages. It is comprised of over 425K aligned clips (763 hours) consisting of both video and audio drawn from over 190 properties covering hundreds of themes and genres. Anim400K is further augmented with metadata including genres, themes, show-ratings, character profiles, and… See the full description on the dataset page: https://huggingface.co/datasets/davidchan/anim400k.