Sign in

Charlie O'Neill

@charlesoneill.bsky.social
11 followers 31 following 0 posts
PostsRepliesMedia
Reposted by Charlie O'Neill
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 23/11/2024
Kind of a broken record here but proceedings.neurips.cc/paper_files/... is totally fascinating in that it postulates two underlying, measurable structures that you can use to assess if RL will be easy or hard in an environment
e introduce the effective horizon, a property of
MDPs that controls how difficult RL is. Our analysis is mo-
tivated by Greedy Over Random Policy (GORP), a simple
Monte Carlo planning algorithm (left) that exhaustively ex-
plores action sequences of length k and then uses m random
rollouts to evaluate each leaf node. The effective horizon
combines both k and m into a single measure. We prove
sample complexity bounds based on the effective horizon that
correlate closely with the real performance of PPO, a deep
RL algorithm, on our BRIDGE dataset of 155 deterministic
MDPs (right).
815029