The dataset includes 100 Reddit posts scraped from healthcare-related subreddits split into 80 labeled training posts, 10 labeled validation posts, and 10 unlabeled test posts. They are classified either as HIPAA violations (yes) or not HIPAA violations (no).
Dataset Details
Dataset Description
A full unlabeled dataset was scraped from nursing, medicine, doctor, physician assistant, CounselingPsychology, and nursepractictioner… See the full description on the dataset page: https://huggingface.co/datasets/ARI-HIPA-AI-Team/Dataset.