A reshaped and edit-distance-annotated version of the
UWV/Leesplank_NL_wikipedia_simplifications_preprocessed
dataset, prepared as the first checkpoint of a Dutch Fluency LoRA project targeting
IBM Granite 4.0 3B Dense.
The original dataset contains ~2.7M paragraph pairs: a Dutch Wikipedia source paragraph
(prompt) and a synthetically simplified version (result). This checkpoint reshapes
those pairs into one text per row, labels… See the full description on the dataset page:
https://huggingface.co/datasets/MichielBuisman/Leesplank-vloeiend-nl-curriculum.