Dataset Card for English-Serbian Semantic Text Similarity Benchmark
Dataset Summary
This dataset is a parallel English-Serbian Semantic Text Similarity (STS) benchmark. It was created to evaluate multilingual English-Serbian language models, with a focus on SBERT (Sentence-BERT) knowledge distillation. The dataset consists of sentence pairs in English and Serbian, along with their semantic similarity scores.
The dataset uses the test split from the original STS benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/STS_parallel_en_sr.