A2C Agent playing PandaReachDense

This is a trained model of an A2C agent playing PandaReachDense using stable-baselines3 and panda-gym.

  • Mean Reward: -0.45 +/- 0.12
  • Result (mean - std): -0.57
Downloads last month
12
Video Preview
loading

Evaluation results