An automatically aligned parallel corpus of well-translated Belgian texts in Dutch and French. The corpus contains texts about Belgium and the Belgian justice system, with over 100.000 tokens.
Part of synthetic multilingual dataset of webcontent from municipalities all over Europe. This dataset was produced within the CEFAT4Cities project. The data is scraped and… See the full description on the dataset page:
https://huggingface.co/datasets/FrancophonIA/public_services_Europe.