Evaluation data used by the AutoSkill skill-discovery loop (see
autoskill_pipeline). Two
files, both multiple-choice video QA:
dev300.json
300
Development set (D_dev), stratified 100/100/100 across short/medium/long duration. Used to drive every discovery-loop cycle (Stage 1 of the pipeline).
pool3000.json
3000
Larger labelled source pool that dev300.json was sampled from (baseline accuracy in an intermediate band, ~60%, at the backbone's… See the full description on the dataset page:
https://huggingface.co/datasets/Cade921/AutoSkill_dev.