GVA_TRANSLATION Dataset is a parallel dataset for machine translation between Valencian (VA) and Spanish (ES).The data in this dataset was shared by the "Dirección General de Política Lingüística de la Generalitat Valenciana" exclusively for use as training data for the ALIA family of models.
The dataset is intended for research in machine translation, cross-lingual NLP, and linguistic analysis.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/gva_translation.