MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page:
https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.