tamil_stories is an open source dataset of instruct-style records generated by scraping publicly available short stories on the following websites.
Apart from scraping and automated cleaning, the data was also tagged manually by a group of volunteers.
This dataset created as part of Aya Open Science Initiative by Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.… See the full description on the dataset page:
https://huggingface.co/datasets/Kaviin/tamil_stories.