Inspired by GLiNER models and its used datasets, we present a Gemini-powered NER Dataset for Bavarian.
The dataset currently features 116,075 sentences from Bavarian Wikipedia, where named entities are found using Gemini 2.0 Flash.
03.07.2025: Initial version of the dataset and public release.
Thankfully, the GLiNER-X community shared their prompt for generating datasets that were used for training the… See the full description on the dataset page:
https://huggingface.co/datasets/bavarian-nlp/gemini-bavarian-ner-v0.1.