AgentVLN adopts a VLM-as-Brain paradigm that decouples high-level semantic reasoning from low-level perception and planning via a plug-and-play skill library. To bridge the gap between strong 2D semantic capabilities of VLMs and the complexities of 3D physical environments, AgentVLN-Instruct tightly aligns high-level instructions with low-level skill invocations.
When generating the dataset using LMDB format, the output structure… See the full description on the dataset page:
https://huggingface.co/datasets/allenxinn/AgentVLN-Instruct.