This model serves as the baseline for the
Ocean Plastic Collection environment, trained and tested on task
1 using the Proximal Policy Optimization (PPO) algorithm.
Environment:
Ocean Plastic Collection
Task:
1
Algorithm:
PPO
Episode Length:
5000
Training
max_steps:
3000000
Testing
max_steps:
150000
Train & Test
Scripts
Download the
Environment