Dataset Card for French Wikipedia Text Corpus
Dataset Description
The French Wikipedia Text Corpus is a comprehensive dataset derived from French Wikipedia articles. It is specifically designed for training language models (LLMs). The dataset contains the text of paragraphs from Wikipedia articles, with sections, footnotes, and titles removed to provide a clean and continuous text stream.
Dataset Details
Features
text: A single attribute containing the full text of… See the full description on the dataset page: https://huggingface.co/datasets/1ou2/fr_wiki_paragraphs.