This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at
https://creativecommons.org/licenses/by-nc/4.0/.
Parallel dataset from Ghanaian news articles.
This is a 100,000-row stratified sample of the original dataset:
Source:… See the full description on the dataset page:
https://huggingface.co/datasets/ghananlpcommunity/english-twi-sentences-non-nouns-sample-100k.