A balanced, three-class sentiment analysis dataset for the Sindhi language (سنڌي) in the Perso-Arabic script, containing 50,000 labeled sentences. Sindhi is a low-resource language spoken by over 30 million people, mainly in Sindh, Pakistan. This dataset supports training and benchmarking of sentiment classification models for Sindhi NLP.
Language: Sindhi (sd), Perso-Arabic script
Task: Sentence-level… See the full description on the dataset page:
https://huggingface.co/datasets/Sanapalijo/sindhi_sentiment.