This agent was trained using Advantage Actor-Critic (A2C) from Stable-Baselines3, designed to solve the robotic manipulation task PandaReachJointsDense-v3. Training was logged using Weights & Biases.
🌎 Environment
Gym ID:PandaReachJointsDense-v3
Task: Robotic manipulation (reach joints) with dense rewards
🧠 Model Configuration
Algorithm: Advantage Actor-Critic (A2C)
Policy Network: Multi-Input Policy
Learning Rate:0.0003
Gamma (Discount factor):0.99
Number of steps per update (n_steps):5
Entropy Coefficient (ent_coef):0
Max Gradient Norm:0.5
Total Training Timesteps:500000
📊 Evaluation Details
Evaluation was performed regularly during training (every 1000 timesteps), and the best-performing models were saved according to the evaluation metrics.