This dataset is a large-scale, cleaned, and standardized collection of 35 diverse publicly available mental health conversational datasets, created for the purpose of LLM finetuning. A comprehensive pipeline was developed to automatically download, extract, analyze, and standardize each source into a unified conversational schema.
Total Conversations: 356791
Train Split: 285432 conversations
Test Split: 71359… See the full description on the dataset page:
https://huggingface.co/datasets/TVRRaviteja/MentalHealthTherapy.