Evaluation code for a key-frame selection benchmark that scores a selected
frame directly — without routing it through a question-answering model, and
without requiring it to match a fixed reference set.
204 films · 2,031 frozen comparison chains · 1,970 learned criteria
The project is three repositories:
Repository
Contents
code
you are here
the evaluation package, all eleven baselines, four numbered scripts