TreeCorpus is a comprehensive, structured dataset derived from the latest Wikipedia dumps, specially processed to serve as high-quality training data for conversational AI models. This dataset transforms Wikipedia's encyclopedic knowledge into a format optimized for natural language understanding and generation tasks.
Size: 26.27 GB (26,272,580,250 bytes)
Examples: 2,882,766 articles
Download Size: 13.33 GB (13,326,529,312 bytes)
Language:… See the full description on the dataset page:
https://huggingface.co/datasets/akkiisfrommars/TreeCorpus.