Sign in

Amir Mesbah

@amirmesbah.bsky.social
135 followers 545 following 20 posts

Graduate Student - Interested in RL and its mathematics 👾 > amirhosein-mesbah.github.io

PostsRepliesMedia
Amir Mesbah @amirmesbah.bsky.social · 23/09/2026
Many people are complaining about the quality of reviews in conferences, but I have not seen any endeavors by community to teach this skill to junior researchers (like me). There were talks and blogs in previous years (e.g., cvpr 2020) but I think we need more of them now.
020
Reposted by Amir Mesbah
Amir-massoud Farahmand @sologen.bsky.social · 08/09/2026
Nice collection of historic AI-related papers and movies! - Minsky's PhD thesis, which he talks about RL (I'd heard of this; never seen!) - Turing's paper on AI - Fukushima's convolutional net - Kismet robot
031
Reposted by Amir Mesbah
Julian Togelius @togelius.bsky.social · 08/12/2025
was at an event on AI for science yesterday, a panel discussion here at NeurIPS. The panelists discussed how they plan to replace humans at all levels in the scientific process. So I stood up and protested that what they are doing is evil. Full post: togelius.blogspot.com/2025/12/plea...
togelius.blogspot.com
Please, don't automate science!
I was at an event on AI for science yesterday, a panel discussion here at NeurIPS. The panelists discussed how they plan to replace humans a...
2827067
Reposted by Amir Mesbah
Claire Vernade @claireve.bsky.social · 12/11/2025
📣 #ICML tutorials: We want to know what *you* would like to learn. This year, Adam White and I are calling for nominations of topics and/or presenters. Until December 7th, you can send us your suggestions, and we will use them to shape the program. icml.cc/Conferences/...
icml.cc
ICML 20256 Call For Tutorials
0159
Reposted by Amir Mesbah
Pablo Samuel Castro @pcastr.bsky.social · 28/10/2025
🚨The Formalism-Implementation Gap in RL research🚨 Lots of progress in RL research over last 10 years, but too much performance-driven => overfitting to benchmarks (like the ALE). 1⃣ Let's advance science of RL 2⃣ Let's be explicit about how benchmarks map to formalism 1/X
1445
Reposted by Amir Mesbah
Amir Balef @amirbalef.bsky.social · 06/10/2025
I am happy to share that our paper "Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning" has been accepted at NeurIPS 2025! Endless thanks to my amazing co-authors @claireve.bsky.social and @keggensperger.bsky.social 📄 Read it on arXiv: arxiv.org/abs/2505.05226 (1/3)
arxiv.org
Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
The Combined Algorithm Selection and Hyperparameter optimization (CASH) is a challenging resource allocation problem in the field of AutoML. We propose MaxUCB, a max $k$-armed bandit method to trade o...
171
Reposted by Amir Mesbah
Claas Voelcker @cvoelcker.bsky.social · 02/10/2025
cvoelcker.de/blog/2025/re... I finally gave in and made a nice blog post about my most recent paper. This was a surprising amount of work, so please be nice and go read it!
media.tenor.com
a close up of a sad cat with the words pleeeaasse written below it
ALT: a close up of a sad cat with the words pleeeaasse written below it
0297
Reposted by Amir Mesbah
Claas Voelcker @cvoelcker.bsky.social · 02/10/2025
cvoelcker.de/blog/2025/re... Here ya go!
cvoelcker.de
Relative Entropy Pathwise Policy Optimization - Technical Overview | Claas A. Voelcker
A lightweight overview of the new REPPO algorithm
111
Reposted by Amir Mesbah
Amir-massoud Farahmand @sologen.bsky.social · 03/08/2025
What are we talking about when we talk about Dynamic Programming? #ReinforcementLearning
281
Amir Mesbah @amirmesbah.bsky.social · 31/07/2025
What if all mathematicians had great visualization skills, tools, and public notes!
010
Reposted by Amir Mesbah
Claire Vernade @claireve.bsky.social · 16/07/2025
Onno and I will be presenting our poster at # W1005 tomorrow (Wed) morning. He made a great thread about it, come chat with us about POMDP theory :)
0195
Reposted by Amir Mesbah
Amir-massoud Farahmand @sologen.bsky.social · 14/07/2025
I will not be at #ICML2025 this year, but 3 of my PhD students at 🤖 Adage (Adaptive Agents Lab) 🤖 are, presenting 3 papers. ⭐ Avery Ma ⭐ Claas Voelcker (cvoelcker.bsky.social) ⭐ Tyler Kastner Meet them to talk about Model-based RL, Distributional RL, and Jailbreaking LLMs.
152
Reposted by Amir Mesbah
Shahab Bakhtiari @shahabbakht.bsky.social · 12/06/2025
Levine's take on the success of LLMs compared to video models is interesting, but I'll expand on how efforts toward AI could take two different paths, and why I think AI and NeuroAI could take different approaches moving forward. 🧵 🧠🤖 #MLSky
172
Reposted by Amir Mesbah
Hafez Ghaemi @hafezghm.bsky.social · 14/05/2025
Preprint Alert 🚀 Can we simultaneously learn transformation-invariant and transformation-equivariant representations with self-supervised learning? TL;DR Yes! This is possible via simple predictive learning & architectural inductive biases – without extra loss terms and predictors! 🧵 (1/10)
15116
Reposted by Amir Mesbah
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 11/05/2025
cleanrl is amazing (github.com/vwxyzjn/clea...) and its structure makes sense for teaching but an actual research codebase should not inherit this style! you do not want this amount of code duplication
github.com
GitHub - vwxyzjn/cleanrl: High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)
High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG) - vwxyzjn/cleanrl
4322
Reposted by Amir Mesbah
Nathan Lambert @natolambert.bsky.social · 18/04/2025
rlhfbook also available on arxiv for SEO 😀 happy friday arxiv.org/abs/2504.12501
arxiv.org
Reinforcement Learning from Human Feedback
Reinforcement learning from human feedback (RLHF) has become an important technical and storytelling tool to deploy the latest machine learning systems. In this book, we hope to give a gentle…
36913
Reposted by Amir Mesbah
Natasha Jaques @natashajaques.bsky.social · 27/03/2025
Recorded a recent "talk" / rant about RL fine-tuning of LLMs for a guest lecture in Stanford CSE234: youtube.com/watch?v=NTSY.... Covers some of my lab's recent work on personalized RLHF, as well as some mild Schmidhubering about my own early contributions to this space
youtube.com
Reinforcement Learning (RL) for LLMs
YouTube video by Natasha Jaques
55110
Reposted by Amir Mesbah
Jakob Foerster @jfoerst.bsky.social · 20/03/2025
PQN puts Q-learning back on the map and now comes with a blog post + Colab demo! Also, congrats to the team for the spotlight at #ICLR2025
0154
Amir Mesbah @amirmesbah.bsky.social · 20/03/2025
Happy #Nowruz and the beginning of the spring!
000
Reposted by Amir Mesbah
Claire Vernade @claireve.bsky.social · 11/03/2025
I’ve put together a short list of opportunities for early career academics willing to come to Europe: www.cvernade.com/miscellaneou... This mostly covers France and Germany for now but I’m willing to extend it. I build on @ellis.eu resources and my own knowledge of these systems.
cvernade.com
Claire Vernade - European career opportunities
European Academic Career Opportunities in 2025
37526
Reposted by Amir Mesbah
Pablo Samuel Castro @pcastr.bsky.social · 05/03/2025
RL is so back! (well, for some of us, it never really left) awards.acm.org/about/2024-t...
awards.acm.org
Andrew Barto and Richard Sutton are the recipients of the 2024 ACM A.M. Turing Award for developing the conceptual and algorithmic foundations of reinforcement learning.
Andrew Barto and Richard Sutton as the recipients of the 2024 ACM A.M. Turing Award for developing the conceptual and algorithmic foundations of reinforcement learning. In a series of papers beginning...
17212
Reposted by Amir Mesbah
Nathan Lambert @natolambert.bsky.social · 26/02/2025
First 11 chapters of RLHF Book have v0 draft done. Should be useful now. Next: * Crafting more blog content into future topics, * DPO+ chapter, * Meeting with publishers to get wheels turning on physical copies, * Cleaning & cohesiveness rlhfbook.com
0489
Reposted by Amir Mesbah
Neuromatch @neuromatch.bsky.social · 24/02/2025
🚨 Neuromatch Academy Course Applications are OPEN for 2025!! 🚨 Get your application in early to be a student or teaching assistant for this year’s courses! Applications are due Sunday, March 23. Apply & learn more: neuromatch.io/courses/ #mlsky #compneurosky #ai #climatesolutions #ScienceEdu 🧪
Applications are now open! 3-week courses: Comp Neuro and Deep Learning. 2-week courses: NeuroAI and Comp Tools for Climate Science.
08674
Reposted by Amir Mesbah
Ben Recht @beenwrekt.bsky.social · 27/01/2025
2014 GoogLeNet: The best image classifier was only trainable using weeks of Google's custom infrastructure. 2018 ResNet: A more accurate model is trainable in a 1/2 hour on a single GPU. What stops this from happening for LLMs?
3519
Reposted by Amir Mesbah
Glen Berseth @glenberseth.bsky.social · 19/01/2025
I am teaching a class on #FoundationalModels for #robotics and Scaling #DeepRL algorithms. This class expands on last year's class and my generalist robotics policies tutorial and code. I plan to share the lectures and code assignments. Starting with the first lectures below.
1216
Amir Mesbah @amirmesbah.bsky.social · 18/01/2025
I wonder why ML conferences insist on uploading workshop videos on SlideShare while they can use YouTube and the benefits of monetization. Talks on SlideShare are really hard to track!
030
Reposted by Amir Mesbah
Pablo Samuel Castro @pcastr.bsky.social · 06/12/2024
i was recently asked to provide 4 "desert island" RL papers. if i were stuck on a desert island i'd hope to have something better to read than #RL papers... but anyway, here's a thread with my choices, maybe you can read them on your flight to @neuripsconf.bsky.social #NeurIPS2024 . Enjoy!
411314
Reposted by Amir Mesbah
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 09/11/2024
If you're an RL researcher or RL adjacent, pipe up to make sure I've added you here! go.bsky.app/3WPHcHg
527127
Reposted by Amir Mesbah
Dylan Foster 🐢 @djfoster.bsky.social · 21/11/2024
As my first post on this platform, allow me to advertise the RL theory lecture notes I have been developing with Sasha Rakhlin: arxiv.org/abs/2312.16730 (shameless repost of my pinned tweet)
421135
Reposted by Amir Mesbah
Marc Lanctot @sharky6000.bsky.social · 21/11/2024
I have become a fan of the game-theoretic approaches to RLHF, so here are two more papers in that category! (with one more tomorrow 😅) 1. Self-Play Preference Optimization (SPO). 2. Direct Nash Optimization (DNO). 🧵 1/3.
2739
Amir Mesbah @amirmesbah.bsky.social · 18/11/2024
Hey academic Bluesky 👀👋
000