The Text parts classification German dataset is a dataset to train language models to classify text parts in scientific papers.
The dataset was created by importing around 10 conference papers and 10 theses and labelling the text parts based on the categories.
The data set was then augmented using synthetic data. For this purpose, the text parts were duplicated using data augmentation.
OpenThesaurus was then used to replace all… See the full description on the dataset page:
https://huggingface.co/datasets/samirmsallem/text_parts_de.