This is a
BGE-M3 model post-trained on the Russian dataset from MMARCO/v2.
The queries are transliterated Russian to English using
uroman.
The model was used for the SIGIR 2025 Short paper: Lost in Transliteration: Bridging the Script Gap in Neural IR.