This dataset contains parsed and heavily cleaned Telegram chat histories, specifically formatted for Causal Language Modeling (CLM) fine-tuning of Base LLMs. It is designed to be used without Chat Templates, allowing the model to learn the natural flow of human conversation.
Total Words (Approx.)
193,555,002… See the full description on the dataset page:
https://huggingface.co/datasets/qzeaq/telegram-dialogues.