PADBen Task 1 is a binary classification dataset for distinguishing between human-authored and LLM-generated paraphrases. This task evaluates whether AI detectors can identify the source of paraphrased text without additional context.
Key Features
Task Type: Binary text classification
Total Samples: 16,233 sentences
Train Split: 12,986 samples (80%)
Test Split: 3,247⦠See the full description on the dataset page: https://huggingface.co/datasets/JonathanZha/PADBen-Task1.