A multimodal dataset for fine-tuning Vision-Language Models (VLMs). It processes egocentric video (Ego4D, EPIC-Kitchens) into structured causal plans, generates 462K multimodal QA pairs across 24 task types, and exports them for LoRA SFT of Qwen3-VL-8B-Instruct.