Dataset Card for ChocoChichia/iliauni-enka-corpus
Overview
This dataset is a Georgian ↔ English parallel corpus scraped from the Ilia State University ENKA Corpus, a large collection of translated texts.
In this dataset, Georgian text is a translation of the provided English sentence.
The corpus is divided into three thematic splits:
Split
Rows
Description
literature
269,972
Literary works, fiction, and translated books.