Continued, off-premise, pre-training of
MedRoBERTa.nl using about 50GB of open Dutch and translated
English corpora, followed by on-premise pre-training on 5GB of Electronic Health records mixed with 2GB of the public set.
All translated (if not with DeepL) with a combination of GeminiFlash 1.5/2.0/GPT4o mini, MariaNMT, NLLB200.
This work was done together with the Amsterdam UMC, in the context of the
DataTools4Heart project.
We were happy to be able to use the
Google TPU research cloud for training the model.