ConstAlignKR is a dataset designed for evaluating and fine-tuning text embedding models, focusing on keyword co-occurrence and alignment with gold-standard references. It is particularly useful for tasks involving semantic textual similarity, keyword extraction, and alignment in Korean text.
License: CC BY 4.0
Languages: Korean
Task: Keyword co-occurrence alignment and semantic similarity evaluation
The dataset… See the full description on the dataset page:
https://huggingface.co/datasets/woojinj-01/ConstAlignKR.