249,949 Azerbaijani web documents annotated with a quality score 0-3.
Used to train a document-level quality classifier for filtering a web
corpus before language-model pretraining.
Texts: sampled from LocalDoc/community_oscar_azerbaijani,
an OSCAR-derived Common Crawl corpus. The texts are NOT original to this dataset.
Labels: generated by the LLM Mistral-Small-24B-Instruct-2501, not by humans.… See the full description on the dataset page:
https://huggingface.co/datasets/LocalDoc/azerbaijani-text-quality-labeled.