This dataset contains multi-turn, persona-driven, code-mixed conversations
generated from real news articles, across 9 Indian
languages, in two script variants:
Native — conversations written in the language's native script,
code-mixed with Romanized English words.
Romanized — conversations fully Romanized (Latin script), code-mixed
with English.
Each language has its own config, loadable independently, e.g.:
from… See the full description on the dataset page:
https://huggingface.co/datasets/LingoIITGN/IndicTalk.