Processed instruction-tuning dataset derived from techiaith/cofnodycynulliad_en-cy, a Welsh–English parallel translation memory published by the Bangor University Language Technologies Unit (Techiaith). Formatted for supervised fine-tuning (SFT) of language models.
Raw pairs
104,738… See the full description on the dataset page:
https://huggingface.co/datasets/locailabs/cofnodycynulliad_en_cy.