Synthetic English audio dataset for time-aware speech understanding, covering temporal localization, temporal description, and timed summaries.
We use 14k hours of audio from YODAS2 English shards, selected from a 24k-hour source pool after language- and silence-ratio filtering. Synthetic annotations were generated for three time-grounded tasks, then filtered through LLM-based verification, deterministic validity checks, and… See the full description on the dataset page:
https://huggingface.co/datasets/ai-sage/TimeGround-1M.