This model is a finetuned version of [Unsloth’s Llama-3.2-11B-Vision-Instruct], trained on a structured cooking QA dataset with three difficulty levels:
Easy – basic cooking questions
Medium – moderately detailed multi-step reasoning
HardQA – complex multi-frame reasoning with deeper analysis
The goal of this model is to analyze cooking videos and answer structured questions about them, focusing on Indian biryani preparation.