This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity.
Subsequently, the selected sentences were translated from Spanish… See the full description on the dataset page:
https://huggingface.co/datasets/gplsi/CA-VA_alignment_test.