For a better dataset description, please visit this GitHub repository prepared by the authors of the article: LINK
This dataset was prepared by converting this dataset Specifically, I've changed the values of scores. If the score was equal or above 2.5, then it was assigned with label "1" (similar). In other cases, label "-1" was assigned (not similar).
I've made these changes in order to apply… See the full description on the dataset page:
https://huggingface.co/datasets/dkoterwa/kor-sts-cosine-embedding-loss.