ShamNER is a curated corpus of Levantine‑Arabic sentences annotated for Named Entities, plus dual annotation to check for consisetency (agreement) across human annotators.
Rounds : pilot, round1–round5 (manual, as a rule quality improved across rounds) and round6 (synthetic, post‑edited). The sythentic data is done by sampling label-rich annotated spans from an MSA project and writing it with an LLM while… See the full description on the dataset page:
https://huggingface.co/datasets/HebArabNlpProject/ShamNER.