The first Old English to Modern English dataset for language translation!
The dataset is structured as source_language, target_language pairs. For more information on how the dataset was extracted, structured, annotated, parsed, etc, you can check this introductory notebook where the data collection methodology is explained and an LLM Gemma-2 model is finetuned on it.
Next you can find the original sources of these texts:
The raw translations come from the excellent work of Dr. Ophelia… See the full description on the dataset page:
https://huggingface.co/datasets/apssg96/the-old-english-dataset.