SID Classification is a Persian (Farsi) dataset designed for the Classification task, specifically targeting document classification of academic articles. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was constructed by collecting academic texts from SID (Scientific Information Database – sid.ir), one of Iran’s major platforms for scientific publications. Each document—formed by concatenating an article’s title and abstract—is… See the full description on the dataset page:
https://huggingface.co/datasets/MCINext/sid-classification.