This dataset consists of synthetic Dutch data, in multiple styles/augmentation methods, categorized by the "type" row, this data has been filtered using Kalamazooter/DutchDatasetCleaner_Bertje.
The main motivation for creating this dataset is the lack of high-quality Dutch datasets, and the fact that existing Dutch datasets have a much smaller amount of code included compared to their English/Multilingual counterparts.
The dataset could be used… See the full description on the dataset page:
https://huggingface.co/datasets/Kalamazooter/GeminiPhiDutch.