MSMARCO-Fa is a Persian (Farsi) dataset created for the Retrieval task, particularly focusing on web search and document ranking. It is a translated version of the original English MS MARCO (Microsoft MAchine Reading COmprehension) dataset and is a key part of the FaMTEB (Farsi Massive Text Embedding Benchmark), under the BEIR-Fa collection.
Language(s): Persian (Farsi)
Task(s): Retrieval (Web Search, Document Ranking)
Source: Translated from the English MS… See the full description on the dataset page:
https://huggingface.co/datasets/MCINext/msmarco-fa.