Embodied-R1 is a 3B vision-language model (VLM) for general robotic manipulation.
It introduces a Pointing mechanism and uses Reinforced Fine-tuning (RFT) to bridge perception and action, with strong zero-shot generalization in embodied… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1-Dataset.