753 book segments of 150-500 characters manually labeled for evaluation of visual descriptiveness, each by two annotators. Segments were sampled from 400 English Project Gutenberg books.
249 segments (those with non-null dataset) were sampled from
https://huggingface.co/datasets/Terraa/vis-desc-large to balance the train set (labels 1, 4, 5). To comply with licenses of texts from that dataset, this one copies its licence and citations below.
=> this yields 1002 rows in total, split into a… See the full description on the dataset page:
https://huggingface.co/datasets/Terraa/vis-desc-small.