Composed of a corpus and software tools, Scientext is a textual database allowing to carry out a linguistic study of the positioning and the reasoning in the scientific writings, thanks to linguistic and statistical queries. Scientext has several sub-corpora: English Corpus of Scientific Writings, French corpus of scientific writings, Communications evaluation body, Corpus of… See the full description on the dataset page:
https://huggingface.co/datasets/FrancophonIA/Scientext.