CQADupstack-unix-Fa is a Persian (Farsi) dataset designed for the Retrieval task, with a focus on duplicate question retrieval. It is a translated version of the "unix" (Unix & Linux Stack Exchange) subforum from the original English CQADupstack dataset, used in the BEIR benchmark, and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection.
Language(s): Persian (Farsi)
Task(s): Retrieval (Duplicate Question Retrieval)
Source:… See the full description on the dataset page:
https://huggingface.co/datasets/MCINext/cqadupstack-unix-fa.