Bilingual (English / German) synthetic corpus of first-person mental-health journal entries paired with structured knowledge graphs. 47,714 samples across 3,410 participants; 41,315 accepted after LLM-based verification.
dataset.jsonl
47,714
full corpus (accept + review + reject), one JSON object per participant-day
graphs.jsonl
48,104
intermediate graphs the pipeline drew from (generator input pool)… See the full description on the dataset page:
https://huggingface.co/datasets/Niklas1102/mentalkg.