A large-scale Twi language dataset containing clean and naturally sounding Twi text across
four distinct styles — monologue, narrative, dialogue, and storytelling — generated
from real Ghanaian news topics as inspiration to keep it grounded on Ghanaian named entities and vocabulary.
This dataset was built to support the development of Twi language models, tokenizers,
and other NLP tools for the Akan language family.