A large-scale multilingual Instagram caption dataset designed for Natural Language Processing (NLP), sentiment analysis, text classification, multilingual language modeling, social-media analytics, caption generation, and AI/ML research.
The dataset contains 105,000+ unique Instagram-style captions covering multiple languages, writing styles, emotions, sentiments, and social-media categories commonly used in Indian and multilingual social-media… See the full description on the dataset page:
https://huggingface.co/datasets/devpatel18042004/Indian-captions.