A 338K-document Hindi corpus (~1.4B tokens) for continued pretraining on
cultural knowledge in figurative language, plus a structured dataset of
16,617 Hindi proverbs (लोकोक्तियाँ) with meanings, recovered via
OCR-repair from a classic proverb dictionary. Each corpus document is natural
Hindi text containing at least one proverb (matched including common surface
variants), with an appended knowledge block listing every… See the full description on the dataset page:
https://huggingface.co/datasets/jiviteshjn/hi-proverbs-cpt.