Paper: Teaching Human Behavior Improves Content Understanding Abilities of VLMs
Website:
https://behavior-in-the-wild.github.io/behavior-llava.html
BLIFT (Behavior-LLaVA Instruction Fine-Tuning) is a large-scale multimodal instruction tuning dataset designed to teach Vision-Language Models (VLMs) human behavior. It contains over 730k images and videos collected from Reddit and YouTube… See the full description on the dataset page:
https://huggingface.co/datasets/behavior-in-the-wild/BLIFT.