ALIA Synthetic MT V2 is a parallel corpus derived from Berria news articles and legal/administrative domains BOPV and Parlamento, comprising content published in 2025 as well as archived material from 2023.
The dataset provides synthetic translations into English and Spanish, generated using two distinct Large Language Models: Qwen3.5-27B and LatxaQ.
This dataset utilizes the following models for translation generation:… See the full description on the dataset page:
https://huggingface.co/datasets/HiTZ/ALIA_syntethic_MT_V2.