This dataset contains Georgian language text pairs designed for natural language processing tasks. Each entry consists of two related Georgian text segments: a shorter "positive" text and a longer "anchor" text that provides additional context or elaboration.
Data Sources
The data for this dataset was gathered from multiple sources: