Dataset concatenating QA datasets with context available in French and open-source.In addition, an augmented version of these datasets has been added (same context but different questions to create data in SQuAD 2.0 format).In total, there are 221,348 training data, 910 validation data and 6,376 test data.In practice, due to the restrictive license for the FQUAD 1.0 dataset, we can only share 200,617 rows of the 221,348 training data and 3,188 rows of the 6,376… See the full description on the dataset page:
https://huggingface.co/datasets/CATIE-AQ/frenchQA.