Munch Hashed Index - Lightweight Audio Reference Dataset
๐ Overview
Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling:
โ Fast duplicate detection across 4.17 million audio samples
โ Efficient dataset exploration without downloading terabytes
โ Quick metadata queries (voiceโฆ See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.