This dataset is a deeply cleaned and structured version of the isek-ai/danbooru-wiki-2024 (Train Split) dataset.
The original Danbooru Wiki contains a massive amount of noise, including raw URLs, complex Textile markup, Wiki syntax, and redundant/malformed aliases. This linguistic noise significantly degrades the performance of embedding models and confuses Large Language Models (LLMs).
To solve this, I filtered out low-information tags and engineered two specialized text fields tailored for… See the full description on the dataset page:
https://huggingface.co/datasets/patvessel/danbooru-rag-v3.