A synthetic dataset for training NER models to detect
and redact HIPAA-sensitive PII from telehealth transcripts.
Custom built dataset with 1600 labeled sentences covering
real-world telehealth scenarios. Created because real
patient data is protected under HIPAA and cannot be
shared publicly.