Sign in

Mattie Fellows

@mattieml.bsky.social
1.7K followers 110 following 21 posts

Reinforcement Learning Postdoc at FLAIR, University of Oxford @universityofoxford.bsky.social All opinions are my own.

PostsRepliesMedia
Mattie Fellows @mattieml.bsky.social · 23/01/2026
www.youtube.com/watch?v=bF_a...
youtube.com
How We Make Hope Normal Again- Green Party Political Broadcast
YouTube video by Green Party of England & Wales
010
Mattie Fellows @mattieml.bsky.social · 21/11/2025
Interested in backprop-free evolution suitable for modern billion parameter models at large population sizes? Super excited to share our latest work, and big congrats to my collaborators @bidiptas13.bsky.social, @juanduquevan.bsky.social and the rest of the @flair-ox.bsky.social team :)
080
Reposted by Mattie Fellows
bidiptas13.bsky.social @bidiptas13.bsky.social · 21/11/2025
Introducing 🥚EGGROLL 🥚(Evolution Guided General Optimization via Low-rank Learning)! 🚀 Scaling backprop-free Evolution Strategies (ES) for billion-parameter models at large population sizes ⚡100x Training Throughput 🎯Fast Convergence 🔢Pure Int8 Pretraining of RNN LLMs
1268
Mattie Fellows @mattieml.bsky.social · 08/09/2025
FLAIR WINTER/SPRING INTERNSHIP! We're looking for two exceptional students to join us on research projects in Oxford from January! Please share with anyone who would be interested. Details below :)
foersterlab.com
Internship - Winter/Spring 2026
We are looking for two talented students to join us for an internship working in FLAIR for 6 months. Students will get the chance to work on current FLAIR projects at the University of Oxford, gaining...
021
Reposted by Mattie Fellows
Pablo Samuel Castro @pcastr.bsky.social · 05/06/2025
PQN, a recently introduced value-based method (bsky.app/profile/matt...) has a similar data-collection as PPO. Although we see a similar trend as with PPO, but much less pronounced. It is possible our findings are more correlated with policy-based methods. 9/
121
Mattie Fellows @mattieml.bsky.social · 30/05/2025
1/2 Offline RL has always bothered me. It promises that by exploiting offline data, an agent can learn to behave near-optimally once deployed. In real life, it breaks this promise, requiring large amount of online samples for tuning and has no guarantees of behaving safely to achieve desired goals.
173
Mattie Fellows @mattieml.bsky.social · 14/05/2025
If you're struggling with the bs Overleaf outage, you can try going to: www.overleaf.com/project/[PROJECTID]/download/zip. to download the zip. It seems to sometimes work after a few minutes
142
Mattie Fellows @mattieml.bsky.social · 25/04/2025
Excited to be presenting our spotlight ICLR paper Simplifying Deep Temporal Difference Learning today! Join us in Hall 3 + Hall 2B Poster #123 from 3pm :)
arxiv.org
071
Reposted by Mattie Fellows
Jakob Foerster @jfoerst.bsky.social · 20/03/2025
PQN puts Q-learning back on the map and now comes with a blog post + Colab demo! Also, congrats to the team for the spotlight at #ICLR2025
0154
Mattie Fellows @mattieml.bsky.social · 20/03/2025
PQN blog 3/3 👉take a look at Matteo's 5-minute blog covering PQN’s key features, plus a Colab demo with JAX & PyTorch implementations mttga.github.io/posts/pqn/ 🔎 For a deeper dive into the theory: blog.foersterlab.com/fixing-td-pa... blog.foersterlab.com/fixing-td-pa... See you in Singapore! 🇸🇬
mttga.github.io
Simplifying Deep Temporal Difference Learning
A modern implementation of Deep Q-Network without target networks and replay buffers.
091
Mattie Fellows @mattieml.bsky.social · 20/03/2025
PQN Blog 2/3: In this blog we show how to overcome `deadly triad' and stabilise TD using regularisation techniques such as LayerNorm and/or l_2 regularisation, deriving a provably stable deep Q learning update WITHOUT ANY REPLAY BUFFER OR TARGET NETWORKS @jfoerst.bsky.social @flair-ox.bsky.social
blog.foersterlab.com
Fixing TD Pt II: Overcoming the Deadly Triad
042
Reposted by Mattie Fellows
Daniel Ansari 🇨🇦 @numcog.bsky.social · 19/03/2025
Are academic conferences in the US a thing of the past?
63013
Mattie Fellows @mattieml.bsky.social · 19/03/2025
PQN Blog 1/3: TD methods are the bread and butter of RL, yet can have convergence issues when used in practice. This has always annoyed me. Find out below why TD is so unstable and how can we understand this instability better using the TD Jacobian. @flair-ox.bsky.social @jfoerst.bsky.social
blog.foersterlab.com
Fixing TD Pt I: Why is Temporal Difference Learning so Unstable?
3193
Mattie Fellows @mattieml.bsky.social · 18/03/2025
Super excited to share our paper, Simplifying Deep Temporal Difference Learning has been accepted as a spotlight at ICLR! My fab collaborator Matteo Gallici and I have written a three part blog on the work, so stay tuned for that! :) @flair-ox.bsky.social arxiv.org/pdf/2407.04811
arxiv.org
3194
Reposted by Mattie Fellows
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 09/11/2024
If you're an RL researcher or RL adjacent, pipe up to make sure I've added you here! go.bsky.app/3WPHcHg
527127