From the Frontier Research Team at Takara.ai we present MovieStills_Captioned_SmolVLM, a dataset of 75,000 movie stills with high-quality synthetic captions generated using SmolVLM.
This dataset contains 75,000 movie stills, each paired with a high-quality synthetic caption. It was generated using the HuggingFaceTB/SmolVLM-256M-Instruct model, designed for instruction-tuned multimodal tasks. The dataset aims to support image captioning tasks, particularly for… See the full description on the dataset page:
https://huggingface.co/datasets/takara-ai/MovieStills_Captioned_SmolVLM.