The camoes_SI dataset is a curated combination of two European
Portuguese sociolinguistic corpora --- Fala Bracarense and
Português Fundamental --- merged into a unified test-only
dataset for evaluating Automatic Speech Recognition (ASR) systems.
All audio is provided as 16 kHz PCM waveforms, accompanied by
speaker metadata and reference transcripts.
This dataset corresponds to the Sociolinguistic Interviews (SI)
category of the CAMÕES… See the full description on the dataset page: https://huggingface.co/datasets/inesc-id/camoes_SI.