CQADupstack-english-Fa is a Persian (Farsi) dataset developed for the Retrieval task, focused on duplicate question detection in community question-answering (CQA) forums. This dataset is a translated version of the "english" (English Language & Usage) StackExchange subforum from the original English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite.
Language(s): Persian (Farsi)
Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page:
https://huggingface.co/datasets/MCINext/cqadupstack-english-fa.