This is a
BGE-M3 model post-trained on the Chinese dataset from MMARCO/v2.
The queries are transliterated Chinese to English using
uroman.
The model was used for the SIGIR 2025 Short paper: Lost in Transliteration: Bridging the Script Gap in Neural IR.