Part 3 adds state-dependent policies and delayed rewards using MENACE/tic-tac-toe.
peterroelants.github.io/posts/post_0...
peterroelants.github.io
From A/B to RL (3/3): Continuous Learning to Delayed Rewards
Use MENACE and tic-tac-toe to move from one-step bandit feedback to state-dependent policies and delayed rewards.