Trained MedRoBERTa-based classifier (full weight update, standard head) on 100.000 OSCAR documents, labeled with GPT4.1-mini as medical/non-medical. Applied to Fineweb2 Dutch.
I am refining this corpus again, with a keyword-exclusion list and a higher model proba, for a version 2.
This is part of the DT4H project (git, website).
If you use this data for your work please use the following citation.
@misc{vanes2026languagecorporadutchmedical,
title={Language… See the full description on the dataset page:
https://huggingface.co/datasets/UMCU/fineweb2medical.nl.