The Expanded Amharic News Dataset is a large-scale, ethically collected corpus of Amharic-language news articles written in Geʽez (Fidel) script, designed to support research in Natural Language Processing (NLP).
This dataset builds upon the “An Amharic News Text Classification Dataset” developed by Israel Abebe Azime and Nebil Mohammed (arXiv link), which categorized Amharic news articles into multiple topical… See the full description on the dataset page: https://huggingface.co/datasets/dagn/expanded-amharic-news-dataset.