This repository contains a fine-tuned RuBERT model optimized for the Semantic Textual Similarity (STS) task on Russian text. It is structured as a Cross-Encoder (Sequence Classification) architecture, passing sentence pairs simultaneously through the network to leverage full cross-attention mechanics.
Framework: PyTorch & Hugging Face Transformers (v5+)
The model evaluates the semantic closeness of two sentences and scores them. Thanks to the Cross-Encoder setup, it excels at capturing nuanced differences (such as negation particles like “не”) and identifying synonyms even when the phrases share zero overlapping words.
Evaluation Results (Test Split)
The model was evaluated on an unseen out-of-domain test split, achieving robust industrial-grade metrics: