This dataset is a high-quality collection of Khmer text images "in the wild." It is a derivative work specifically enhanced for Visual Grounding and High-Accuracy OCR fine-tuning.
Original Source: Khmer Word Dataset (Kaggle)
Original Owner: Saly KEO
Modification: The original dataset (which provided single text labels per image) has been re-annotated and processed to include normalized bounding… See the full description on the dataset page:
https://huggingface.co/datasets/vichetkao/wild_khmer.