Unified Spanish Misinformation and Satire Corpus (USMSC)
Dataset Description
This dataset is a comprehensive, deduplicated, and systematically structured corpus for domain-specific misinformation detection in Spanish social media text. It addresses the critical gap in Spanish-language resources by unifying multiple distinct datasets into a single, highly refined corpus.
Crucially, this dataset employs a three-class formulation (Fake, Real, Satire). Recent… See the full description on the dataset page: https://huggingface.co/datasets/gabrielhuav/Unified-and-Balanced-Spanish-Fake-News-Corpus.