This repository contains the refreshed unified TELL binary-provenance corpus for AI-generated text detection. A subset of this was used in the paper: arxiv.org/abs/2605.27921
Dataset credit: Aldan Creo and Suraj Ranganath.
Rows: 9,179,122
Label convention: 0 = human, 1 = AI
Main data file: data/unified_tell_dataset.parquet
Metadata/counts: dataset_summary.json
Schema includes both raw lang and normalized… See the full description on the dataset page:
https://huggingface.co/datasets/suraj-ranganath/unified_tell_dataset.