Describe Anything: Detailed Localized Image and Video Captioning
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Dataset Card for DLC-Bench
Dataset for detailed localized captioning benchmark (DLC-Bench).
CC BY-NC-SA 4.0… See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/DLC-Bench.