Model type:
Finedefics is an open-source MLLM that enhances the model's FGVR capability by incorporating informative attribute descriptions of objects into the training phase.
It is an auto-regressive language model, based on the transformer architecture.
Base MLLM:
HuggingFaceM4/idefics2-8b
Idefics2 is licensed under the Apache 2.0 license, and we release the Finedefics checkpoints under the same license.
A collection of 6 fine-grained visual recognition datasets, including Stanford Dog-120, Bird-200, FGVC-Aircraft, Flower-102, Oxford-IIIT Pet-37, and Stanford Car-196.