M1llion-Lang is a high-quality, large-scale multilingual instruction dataset designed for training and fine-tuning large language models (LLMs) to understand and generate text in 20+ languages with natural emoji expression and cultural nuance.
Total Size: ~4GB (JSON Lines format)
Languages: 20 languages covering 95%+ of global internet usersFormat: Conversational JSONL with⦠See the full description on the dataset page:
https://huggingface.co/datasets/m1llion-ai-high-end-group/m1llion-lang.