This dataset contains a curated selection of sent emails from the historical Enron email corpus. It is designed specifically for text privatization, author identification, stylometry profiling, and text classification tasks, focusing on the most prolific writers in the corpus.
author: The identifier of the email sender (e.g., dasovich-j, germany-c). There are 28 unique authors in total, heavily represented by… See the full description on the dataset page:
https://huggingface.co/datasets/sjmeis/enron28.