Views
No views yet
Reason over video with a video-language model; plus a LoRA fine-tune sketch.
| Base model | Qwen/Qwen2-VL-2B |
| Task | video question answering |
| Training objective | Video-grounded next-token prediction (inference; optional LoRA SFT). |
| Track | LM · Language & multimodal |
| Built on | QwenLM/Qwen2-VL |
| Notebook | |
| Compute / storage / time | GPU required — see the Compute · storage · time table in the notebook |
HfApi().upload_folder(...)) — the checkpoint + metrics.json + figures replace this placeholder.metrics.json · [ ] add figures · [ ] swap in the real results card1@misc{ropedia_academy,
2 title = {Ropedia Academy: an interactive course on embodied & spatial AI},
3 author = {Ropedia Academy},
4 year = {2026},
5 howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
6}