shawnhymel.com
Reinforcement Learning Part 13: Policy Gradient Causality Trick and REINFORCE - Shawn Hymel
In the previous post, we showed how we can substitute our usual ε-greedy policy with a parameterized approximation (often a neural network), we then derived