The dataset is composed of 1000 images containing Bangla and Assamese text.
Bangla and Assamese are closely related languages with sharing the
Bengali–Assamese script and have similar lexical
constructs (
https://en.wikipedia.org/wiki/Bengali%E2%80%93Assamese_script#cite_note-MajR-22).
The text is this dataset has been generated from Open English Bible
(
https://openenglishbible.org/oeb/2022.1/OEB-2022.1-US.txt) by:
Selecting the first 500 lines containg between 50 and… See the full description on the dataset page:
https://huggingface.co/datasets/anrikus/lexical_diff_bangla_assamese_v2.