Small (100k+) sythetic dataset for fine-tuning text embedding models for Ukraininan language (STS task)
If you find this work useful, please consider citing:
@article{SED-UA-small2025,
title = {SED-UA-SMALL: Ukrainian synthetic dataset for text embedding models},
volume = {17},
ISSN = {2663-0001},
doi = {10.23939/sisn2025.17.403},
url =… See the full description on the dataset page:
https://huggingface.co/datasets/suntez13/sed-ua-small-sts-v1.