SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD.
In this repository we share the text-only… See the full description on the dataset page:
https://huggingface.co/datasets/scilons/SciLaD-en-dedup-v1.