A unified etymological knowledge graph built from five publicly available datasets. Contains 22.7 million relations across 6.6 million words in 5,529 languages, covering cognate pairs, derivation chains, borrowing paths, and more.
The goal was to bring together the best open etymology data into a single, queryable SQLite database with a normalized schema — one you can point a graph search at without wrangling five different formats.
What's in the… See the full description on the dataset page: https://huggingface.co/datasets/danielquillanroxas/unified-etymology-graph.