FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models.
More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision
available_subsets =… See the full description on the dataset page:
https://huggingface.co/datasets/HuggingFaceM4/FineVision.