FETV is a benchmark for Fine-grained Evaluation of open-domain Text-to-Video generation
Overview
FETV consist of a diverse set of text prompts, categorized based on three orthogonal aspects: major content, attribute control, and prompt complexity.
Dataset Structure
Data Instances
All FETV data are all available in the file fetv_data.json. Each line is a data instance, which is formatted as:
{
"video_id": "1006807024",
"prompt": "A mountain… See the full description on the dataset page: https://huggingface.co/datasets/lyx97/FETV.