This dataset provides the ToS benchmark annotations used in ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation.
ToS is designed for continual learning in open-ended image-to-text generation. Each image is assigned to a task by its dominant visual topic. Other visible topics remain in the sample, so tasks shift over time while shared concepts can still reappear across tasks.