Conversations: 240,585
File size: ~143 MB
Format: JSON array of chat conversations
Content: synthetic Japanese multi-turn conversations for language tasks
This dataset contains high-quality synthetic Japanese question-answering pairs, generated from various rich linguistic and encyclopedic sources. It is primarily designed to improve language models' understanding of Japanese grammar, vocabulary, idioms… See the full description on the dataset page:
https://huggingface.co/datasets/bunbohue/Japanese-Knowledge-Base-Synthetic-Data.