Dataset Card for escagleu-64K corpus
Dataset Description
Dataset Summary
This is the second version of escagleu-64k, a parallel corpus containing approximately 64k sentences translated across Spanish, Catalan, Valencian Catalan, Galician, and Basque.
The original sentences are in Spanish and are sourced from the Spanish Common Voice Corpus.
This corpus was prepared with the goal of creating a parallel speech dataset for these languages using the Common Voice… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/escagleu-64k.