MDTA is a benchmark for AI-generated text detection. It pairs human-written and LLM-generated answers across five domains, four open-weights models, and three sampling temperatures, and augments each LLM response with three adversarial paraphrases (including constrained letter-avoidance rewrites).
The dataset spans 24,322 prompt-aligned questions (around 642,000 text samples when counting all generated and adversarial responses).… See the full description on the dataset page:
https://huggingface.co/datasets/nsp909/MDTA.