Studio-recorded Filipino read, spontaneous, and word-level speech — 125 speakers, packaged as ready-to-stream Parquet.
This is the Filipino Speech Corpus (Sagum), recorded in a controlled setting and
hand/machine transcribed with Transcriber. This repo
repackages the original .wav + .trs volumes as segment-level Parquet with
inline audio, so you can stream it without… See the full description on the dataset page:
https://huggingface.co/datasets/sapinsapin/filipinospeechcorpus.