Pashto Multimodal v1: A Decade of Pashtun Digital Activism
Dataset Summary
This is a curated, multilingual, and multimodal dataset of ~5,900 high-quality image-text pairs, extracted from the personal X (Twitter) archive of @Pashto_lab. This archive represents over a decade of active digital advocacy (since 2012) by a Pashtun AI researcher and activist based in Japan.
The dataset is a rich, firsthand chronicle of contemporary Pashtun socio-political discourse… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto_multimodal_v1.