One content lemma per surface token, in context. Used to train and evaluate the
Shoshan lemmatizer.
train.csv / dev.csv / test.csv
in-domain (Knesset + Wikipedia, IAHLT UD)
191k / 11k / 11k
ood.csv (+ ood_Bagatz.csv, ood_GeekTime.csv, ood_Dicta.csv)
out-of-domain benchmark, 100 sentences/domain
~5k
Columns: form, lemma, pos, sentence, source, sent_id.
Load with the package:
from… See the full description on the dataset page:
https://huggingface.co/datasets/HebArabNlpProject/shoshan-data.