Fetch from
https://huggingface.co/datasets/HeNLP/HeDC4
Clean with IQR bounds forumula (Bounds: -376.0 <= length <= 1168.0)
Aggressive clean: keep only sentences with Hebrew/basic punctuation
Remove duplicate punctuation
Transform HeDC4-enhanced-v1.csv to new CSV with id,text headers, , separator and " escape char (default)
Add diacritics, stress marks… See the full description on the dataset page:
https://huggingface.co/datasets/Phonikud/heb-text.