Sign in

Peter Rlnts

@peterroelants.bsky.social
33 followers 214 following 12 posts

Carbon-based silicon enthusiast @ anam.ai - peterroelants.github.io

PostsRepliesMedia
Peter Rlnts @peterroelants.bsky.social · 24/06/2026
Part 3 adds state-dependent policies and delayed rewards using MENACE/tic-tac-toe. peterroelants.github.io/posts/post_0...
peterroelants.github.io
From A/B to RL (3/3): Continuous Learning to Delayed Rewards
Use MENACE and tic-tac-toe to move from one-step bandit feedback to state-dependent policies and delayed rewards.
000
Peter Rlnts @peterroelants.bsky.social · 24/06/2026
Part 2 moves from fixed experiments to online learning: multi-armed bandits, probability matching, and Thompson sampling. peterroelants.github.io/posts/post_0...
peterroelants.github.io
From A/B to RL (2/3): Multi-Armed Bandits
Move from fixed A/B testing to online learning with Bayesian multi-armed bandits, probability matching, and Thompson sampling.
110
Peter Rlnts @peterroelants.bsky.social · 24/06/2026
I finally cleaned up some old @jupyter.org notebook drafts from when I was teaching myself #ReinforcementLearning They became From A/B to RL: a 3-part series bridging A/B testing and reinforcement learning, starting with Bayesian A/B testing: peterroelants.github.io/posts/post_0...
peterroelants.github.io
From A/B to RL (1/3): Bayesian A/B Testing
Connect fixed A/B testing to Bayesian posterior uncertainty over click-through rates, and use that uncertainty to make one final decision.
110
Peter Rlnts @peterroelants.bsky.social · 19/05/2026
While initially I spent an absurd amount of time creating the original Bokeh visualizations, its great to be able to adapt them quickly and easily nowadays without having to write the code myself.
000
Peter Rlnts @peterroelants.bsky.social · 19/05/2026
I started picking up some of my old notes (like these) and use LLMs to clean them up and publish them. These notes were probably 80% done, the core aspects were already there. But it still took a lot of hand-holding to get them in the current state using frontier models
100
Peter Rlnts @peterroelants.bsky.social · 19/05/2026
I started writing these blogposts as @jupyter.org notebooks a long while ago (2019 I think) when teaching myself on reinforcement learning and going in-depth into probability theory.
111
Peter Rlnts @peterroelants.bsky.social · 19/05/2026
The second post was prompted by trying to interpret what it means to be a good prior, which led me to learn about Pólya Urns, which I think show parallels with Thompson Sampling popular in multi-armed bandits - peterroelants.github.io/posts/beta-p...
peterroelants.github.io
Beta Priors for Self-Reinforcing Binary Decisions
See how Beta-Bernoulli updates work one observation at a time in A/B tests and bandits, and how the Pólya urn explains self-reinforcing feedback.
110
Peter Rlnts @peterroelants.bsky.social · 19/05/2026
The first post is inspired by Thomas Bayes's canonical work "An Essay towards Solving a Problem in the Doctrine of Chances" - peterroelants.github.io/posts/beta-d...
peterroelants.github.io
The Beta Distribution: A Distribution Over Probabilities
Understand the Beta distribution as a model for unknown success probabilities, and derive the Beta-Bernoulli update for binary data.
110
Peter Rlnts @peterroelants.bsky.social · 19/05/2026
I published 2 new blogposts recently digging into the Beta distribution: - peterroelants.github.io/posts/beta-d... - peterroelants.github.io/posts/beta-p...
100
Peter Rlnts @peterroelants.bsky.social · 01/11/2025
Hopefully it’s useful for anyone exploring flow matching for generative modeling. Writing it certainly helped solidify my own understanding.
000
Peter Rlnts @peterroelants.bsky.social · 01/11/2025
I've been working with flow matching models for video generation for a while, and recently went back to my old notes from when I was first learning about them. I cleaned them up and turned them into this blog post.
100
Peter Rlnts @peterroelants.bsky.social · 01/11/2025
New blogpost: Flow Matching: A visual introduction peterroelants.github.io/posts/flow_m...
100