[Paper]
[Github]
This dataset contains a total of 15,000 pieces of real or AI-generated multimodal news (image-caption pairs) -- a training set of 10,000 pairs, a validation set of 2,500 pairs, and five test sets of 500 pairs each. Four of the test sets are out-of-domain data from unseen news publishers and image generators to evaluate detector's generalization ability.
=== Data Source (News Publisher + Image Generator)… See the full description on the dataset page:
https://huggingface.co/datasets/Gouge666/mirage-news.