The dataset contains 16_958_654 extracted ontologies from a subset of selected wikipedia articles.
The dataset was created via LLM processing a subset of the English Wikipedia 20231101.en dataset.
The initial knowledge base dataset was used as a basis to extract the ontologies from.
Pipeline: Wikipedia article → Chunking → Fact extraction (Knowledge base dataset) → Ontology extraction from facts → Ontologies… See the full description on the dataset page:
https://huggingface.co/datasets/Jotschi/wikipedia_knowledge_graph_en.