This dataset is derived from corporate names and their corresponding kana obtained from the Corporate Number Publication Site. We applied the following processing steps:
After removing corporate designations (e.g., "株式会社"), we extracted only those corporate names composed entirely of English letters.
Removed all corporate type identifiers.
Converted full-width characters to half-width.
Excluded compound words such as those in camelCase.
Converted all text to lowercase.… See the full description on the dataset page:
https://huggingface.co/datasets/m7142yosuke/english2kana-v1.