Pietro Novelli @pienovelli.bsky.social · 09/01/2025For the past four years, I’ve been working on a topic that’s both fascinating and challenging to explain. In this post, I’ve tried to present The Operator Way — a paradigm for understanding dynamical processes — in plain, approachable terms. pietronvll.github.io/the-operator... 170
Pietro Novelli @pienovelli.bsky.social · 12/12/2024By the time I finished working on this paper, I had more research questions than when I started. I take this fertility of ideas as a very good sign 😃. If you’re in Vancouver, consider checking it out. I’ll be at the West Ballroom A-D from 16:30 to 19:30, poster #6907 020
Pietro Novelli @pienovelli.bsky.social · 12/12/2024To add some flesh around this core idea, we developed a neat theoretical foundation that combines conditional mean embeddings and policy mirror descent. This foundation ultimately leads to sample complexity results, highlighting the interplay between exploration and exploitation. 120
Pietro Novelli @pienovelli.bsky.social · 12/12/2024The return is a (conditional) expected value, and we realized that there are now mature ML tools to model such expected values directly, avoiding the solution of intermediate and more difficult problems. 120
Pietro Novelli @pienovelli.bsky.social · 12/12/2024So, what’s all this fuss about? Reinforcement learning, in essence, is an optimization problem: we want to maximize returns. 120
Pietro Novelli @pienovelli.bsky.social · 12/12/2024This quote neatly encapsulates the core of our “Operator World Models for Reinforcement Learning” which we’re presenting today at @NeurIPS. arxiv.org/abs/2406.19861arxiv.orgOperator World Models for Reinforcement LearningPolicy Mirror Descent (PMD) is a powerful and theoretically sound methodology for sequential decision-making. However, it is not directly applicable to Reinforcement Learning (RL) due to the inaccessi... 120
Pietro Novelli @pienovelli.bsky.social · 12/12/2024In his book “The Nature of Statistical Learning” V. Vapnik wrote: “When solving a given problem, try to avoid a more general problem as an intermediate step” 183