🧠 DBpediaOntoTrain: A Quality-Segmented Ontology Dataset for LLM Pretraining
📘 Overview
DBpediaOntoTrain is a dataset of 1,766 OWL ontologies in Turtle format, extracted from DBpedia Archivo and prepared for continual pretraining of Large Language Models (LLMs) in ontology generation and completion tasks.
Each ontology is analyzed using a set of semantic quality metrics, tokenized using the LLaMA 3.2 tokenizer, and sorted by Quality Score (QS). The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/DBpediaOntoTrain.