ViTextVQA là một dataset Visual Question Answering (VQA) dành cho tiếng Việt, tập trung vào khả năng đọc hiểu chữ xuất hiện trong ảnh (scene text), dựa trên bài báo ViTextVQA: A Large-Scale Visual Question Answering Dataset for Evaluating Vietnamese Text Comprehension in Images (ArXiv 2404.10652).Bản release trên Hugging Face này đã được chỉnh sửa cấu trúc và bổ sung các… See the full description on the dataset page: https://huggingface.co/datasets/nhonhoccode/ViTextVQA.