The dataset has 68 character classes, 10 digit classes, and 320 syllable classes.
For constructing the dataset, 1072 Thai native speakers wrote on collection datasheets
that were then digitized using a 300 dpi scanner.
De-skewing, detection box and segmentation algorithms were applied to the raw scans
for image extraction. The dataset, unlike all other known Thai handwriting datasets, retains
existing noise, the white background, and all artifacts generated by scanning.