MIRALL is a paired human/machine corpus for machine-generated text (MGT) detection in Catalan.
Every human anchor from IberAuTexTification (IberLEF 2024)
is paired one-to-one with a 2024-2026 model generation produced from that same anchor, giving a
balanced human vs machine detection task across 5 domains and 4 generators.
7,283 pairs (14,566 texts) | 5 domains | 4 generators | 44 prompt cells | locked 70/10/20… See the full description on the dataset page:
https://huggingface.co/datasets/cescgr17/mirall.