Search 3.2M models and datasets…
⌘K
Chat
Models
Datasets
Deploy
Pricing
Docs
Chat
Models
Datasets
Deploy
More
Marco_Longspeech – Dataset by ATH-MaaS | AlphaNeural AI
Is this your dataset? Claim it with the Hugging Face account that owns it.
ATH-MaaS
/
Marco_Longspeech
like
0
automatic-speech-recognition
audio-classification
text-generation
en
zh
apache-2.0
10K<n<100K
audio
2601.13539
us
audio
speech
asr
speech-recognition
question-answering
summarization
translation
emotion-recognition
speaker-diarization
Views
No views yet
Dataset card
Files and Versions
Community
Use
Use this dataset
Marco-LongSpeech Dataset
Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks designed to benchmark Large Language Models on lengthy audio inputs.
📊 Dataset Statistics Task Statistics
Task Train Val Test Total Unique Audios
ASR 71,275 15,273 15,274 101,822 101,822
Temporal_Relative_QA 5,886 1,261 1,262 8,409 8,409
summary 4,366 935 937 6,238 6,238… See the full description on the dataset page:
https://huggingface.co/datasets/ATH-MaaS/Marco_Longspeech
.