This dataset contains document annotations produced with propella-1-4b, a small multilingual LLM that annotates text documents across six categories: core content, classification, quality & value, audience & purpose, safety & compliance, and geographic relevance. The annotations can be used to filter, select, and curate LLM training data at scale.
Each document is annotated across 18 properties organized into six categories:
Category
Property
Description… See the full description on the dataset page:
https://huggingface.co/datasets/openeurollm/propella-annotations.