EduVQA-Alpha is a multilingual educational dataset designed for video question-answering (VideoQA). It consists of academic videos, annotated with synthetic question-answer (QA) pairs, in English and Persian. Videos are curated to reflect diverse academic topics and teaching styles, supporting multilingual Retrieval-Augmented Generation (RAG) tasks.
The dataset employs CLIP-SSIM Adaptive Chunking for video segmentation… See the full description on the dataset page: https://huggingface.co/datasets/UIAIC/EduViQA.