Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in VisionโLanguageโAction (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issueโฆ See the full description on the dataset page:
https://huggingface.co/datasets/OpenMOSS-Team/OmniAction.