This is a trained model of a REINFORCE (Monte Carlo Policy Gradient) agent playing Pixelcopter-PLE-v0,
implemented from scratch with PyTorch, as part of the Hugging Face Deep RL Course, Unit 4.
Results
Mean reward over 10 evaluation episodes: 46.40 +/- 37.34