A high-quality synthetic Bengali OCR dataset for fine-tuning vision-language models like DeepSeek-OCR 2. Generated using 100+ professional Bengali Unicode fonts and 13K+ unique Bengali words with advanced text rendering via FreeType and HarfBuzz.
Language: Bengali (বাংলা)
Task: Optical Character Recognition (OCR)
Format: Conversation-based (vision-language)
Total Samples: 30,000
Train: 27,007 samples
Validation: 2,993… See the full description on the dataset page:
https://huggingface.co/datasets/rifathridoy/bengali-ocr-synthetic.