Dataset concatenating Simplification datasets, available in French and open-source.There are a total of 785,625 rows, of which 781,801 are for training, 2,385 for validation and 1,439 for testing.
from datasets import load_dataset
dataset = load_dataset("CATIE-AQ/frenchSIMPLIFICATION")
Dataset Original
Splits
Note
clear
4,196 train / 300 validation / 100 test
wikilarge
296,402 train /… See the full description on the dataset page:
https://huggingface.co/datasets/CATIE-AQ/frenchSIMPLIFICATION.