This dataset contains expert trajectories generated by a Proximal Policy Optimization (PPO) reinforcement learning agent trained on 4 environments of the Distracting Control Suite. For each environment we collect data with different levels of distraction, which we define below, and masks for the agent.
Levels of distraction:
None: Vanilla DeepMind Control Suite without visual distractions. The environment uses the default static background… See the full description on the dataset page:
https://huggingface.co/datasets/hamza-adnan/visual_distracting_control_suite.