A dataset for fine-tuning semantic search and text embedding models on Russian classical literature.
284,275 triplets (anchor, paraphrase, negative) extracted from Russian classical literary texts.
Hard negatives were mined using microsoft/harrier-oss-v1-270m.
paraphrase
A short description of the anchor's content (avg ~44… See the full description on the dataset page:
https://huggingface.co/datasets/RafaelUI/RuLitSearch.