High-quality Japanese word embeddings trained on Wikipedia using MeCrab morphological analyzer.
This dataset contains pre-trained Japanese word embeddings optimized for use with MeCrab, a high-performance morphological analyzer.
Key Features:
✅ Trained on Japanese Wikipedia
✅ Zero-copy binary format (MCV1) for fast loading
✅ Compatible with MeCrab Python API
✅ 300-dimensional vectors
✅ ~100,000 vocabulary size… See the full description on the dataset page:
https://huggingface.co/datasets/KitaSan/mecrab-jawiki-word2vec.