Diverse, fluent Dutch pretraining text synthesized from
guust-franssens/belgisch-staatsblad
(CC0, Belgian official-gazette filings). Adds Belgium/Flanders coverage to Dutch LM pretraining
mixes, where clean Belgian-Dutch prose is otherwise scarce.
The source text is noisy OCR from scanned PDFs, but its metadata (company, juridical form,
act type, city, date) is clean. A local LLM (google/gemma-2-9b-it)
"launders" the OCR + metadata… See the full description on the dataset page:
https://huggingface.co/datasets/pdelobelle/staatsblad-synth-nl.