SyosetuNames-3.5M: Japanese Light Novel Character Names Corpus
Overview
This dataset extracts fictional character names from the publicly available text of novels on the Japanese light novel platform "Shōsetsuka ni Narō" (syosetu.com), containing approximately 3.5 million unique original and cleaned names. This dataset aims to provide resources for culturally sensitive Natural Language Processing (NLP) tasks, such as name generation, Named Entity Recognition (NER) model… See the full description on the dataset page: https://huggingface.co/datasets/Sunbread/SyosetuNames-3.5M.