This repository contains 6,493 synthetic Dutch clinical documents with
character-offset de-identification spans. It contains no real patient notes or
personal information. The corpus was used to train meddeid-dutch-synth.
All 6,493 records are exposed together through one conventional Hugging Face
train split. There is no publisher-defined validation split. Here train
means the complete model-development corpus; users… See the full description on the dataset page:
https://huggingface.co/datasets/stighellemans/meddeid-dutch-synthetic-corpus.