Sign in

Theresa Eimer

@theeimer.bsky.social
1.1K followers 480 following 83 posts

RL researcher looking for DACs // What is this AutoRL anyway? she/her Currently: Leibniz Uni Hannover Previously: Uni Freiburg (Master's) | Meta AI London (Intern) Always & Forever: AutoRL.org

PostsRepliesMedia
Reposted by Theresa Eimer
Clément Canonne @ccanonne.github.io · 06/03/2026
Leibniz, looking at the universe: "Why is there something instead of nothing?" Me, looking at my Outlook calendar: same
210117
Theresa Eimer @theeimer.bsky.social · 24/02/2026
A WIP, more or less. But here's the matching data: huggingface.co/datasets/aut...
huggingface.co
autorl-org/arlbench · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Theresa Eimer @theeimer.bsky.social · 23/02/2026
Untunable? Very uncharitable, don't you think? If you look at a DQN hyperparameter eCDF, you clearly see that it can perform well! Just, you know,... incredibly rarely 🤡
eCDF plot von DQN. The curve is pretty bad, 80% of configurations are below 50% max performance.
120
Reposted by Theresa Eimer
ELLIS @ellis.eu · 04/02/2026
Meet @theeimer.bsky.social, Postdoc at @unihannover.bsky.social 🇩🇪, ELLIS Member working on AutoRL, hyperparameter optimization & Reinforcement Learning evaluation. Her challenge at work: bridging #AutoML + RL and sharpening communication to make her work clear to both communities. #WomenInELLIS
092
Theresa Eimer @theeimer.bsky.social · 18/12/2025
Living document means we welcome any discussion and additions! Use the Issues and PRs in the repository to improve this document and hopefully it can be a resource for years to come.
100
Theresa Eimer @theeimer.bsky.social · 18/12/2025
A little Christmas present from and for the COSEAL community: a compilation of the best research practices and workflow recommendations in a living document: github.com/coseal/COSEA... Our goal is to improve research quality in meta-algorithmics and to give new researchers an easier start 💪
github.com
100
Theresa Eimer @theeimer.bsky.social · 11/12/2025
I'm super fascinated by the randomized results in the talk, though. Could be hard to spot, but basically I tuned PPO evaluation either 1 seed per HP config, 20 seeds or 20 runs with random seed, n_envs, hidden size and activation. The latter performed way better on the default eval and in transfer!
010
Theresa Eimer @theeimer.bsky.social · 11/12/2025
...But that's probably a very specific point of view. Seems very difficult to me currently to do evaluations for algorithms that are supposed to solve everything at once if we focus mostly on solution scores or fixed benchmarks.
110
Theresa Eimer @theeimer.bsky.social · 11/12/2025
I just gave a talk (aka thinking out loud) at the BeNRL seminar about expressiveness of evaluations. I landed closer to "show people more and potentially weird things" rather than "standardize the setup"... theeimer.github.io/assets/pdf/s...
theeimer.github.io
251
Theresa Eimer @theeimer.bsky.social · 31/10/2025
Foundation models on the AutoML podcast 2/3: are LLMs killing AutoML? It's probably not that simple. Listen for more details 😉
010
Theresa Eimer @theeimer.bsky.social · 24/10/2025
Stealing all of the recommendations! This made me think of The Left Hand Of Darkness, though I guess that's actually almost the opposite, communication bridging a seemingly impossible gap in understanding each other...
030
Theresa Eimer @theeimer.bsky.social · 22/09/2025
I fell into a hole, but made it out again with new episodes! This is part one of three of an accidental series on foundation models. The next parts will be released in October and November, so stay tuned!
050
Theresa Eimer @theeimer.bsky.social · 28/08/2025
Great opportunity to work with great people. Go apply!
010
Reposted by Theresa Eimer
Julian Togelius @togelius.bsky.social · 13/08/2025
New blog post: AI Allergy. On my increasing disgust with the AI discourse, even though I still like the technical and philosophical. And how I wish I could be excited about AI again. togelius.blogspot.com/2025/08/ai-a...
togelius.blogspot.com
AI Allergy
I remember being excited about AI. I remember 20 years ago, being excited about neuroevolutionary methods for learning adaptive behaviors in...
69018
Reposted by Theresa Eimer
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 11/07/2025
It is time
5566
Reposted by Theresa Eimer
Mark A. Hanson @hansonmark.bsky.social · 10/07/2025
The "reproducibility crisis" in science constantly makes headlines. Repro efforts are often limited. What if you could assess reproducibility of an entire field? That's what @brunolemaitre.bsky.social et al. have done. Fly immunity is highly replicable & offers lessons for #metascience A 🧵 1/n
11318173
Reposted by Theresa Eimer
Antonin Raffin @araffin.bsky.social · 07/07/2025
Need for Speed or: How I Learned to Stop Worrying About Sample Efficiency Part II of my blog series "Getting SAC to Work on a Massive Parallel Simulator" is out! I've included everything I tried that didn't work (and why Jax PPO was different from PyTorch PPO) araffin.github.io/post/tune-sa...
araffin.github.io
Getting SAC to Work on a Massive Parallel Simulator: Tuning for Speed (Part II) | Antonin Raffin | Homepage
This second post details how I tuned the Soft-Actor Critic (SAC) algorithm to learn as fast as PPO in the context of a massively parallel simulator (thousands of robots simulated in parallel).
4358
Reposted by Theresa Eimer
Mattie Fellows @mattieml.bsky.social · 30/05/2025
1/2 Offline RL has always bothered me. It promises that by exploiting offline data, an agent can learn to behave near-optimally once deployed. In real life, it breaks this promise, requiring large amount of online samples for tuning and has no guarantees of behaving safely to achieve desired goals.
173
Theresa Eimer @theeimer.bsky.social · 27/05/2025
Crazy volume! On the other hand, not that surprising. We also got one of these and only did so because it was such a good deal that even if our complete lack of experience makes research on it hard, we can use it for teaching only, and be okay with spending the money. I doubt we're the only ones!
020
Reposted by Theresa Eimer
Katharina Eggensperger @keggensperger.bsky.social · 21/05/2025
📢 Only 3 Weeks to Go! The AutoML summer school (June 10-13th) is just around the corner, and there is not much time left to register! ---> www.automlschool.org <--- 👇 We added several new speakers to the program
automlschool.org
AutoML School 2025
Scope AutoML has become a cornerstone in the toolkit of many developers and researchers. With the rise of foundation models, AutoML's potential has expanded even further, enabling smarter, more powerf...
174
Reposted by Theresa Eimer
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 19/05/2025
Going to the hospital because I broke my wrist smashing the endorse button: www.understandingai.org/p/i-got-fool...
understandingai.org
I got fooled by AI-for-science hype—here's what it taught me
I used AI in my plasma physics research and it didn’t go the way I expected.
611929
Reposted by Theresa Eimer
Rieks op den Akker @rieks123.bsky.social · 14/05/2025
We can only presume to build machines like us once we see ourselves as machines first. Abeba Birhane (2022, p. 13) This is the core. So true.
2299
Reposted by Theresa Eimer
Cathy Wu @cathywu.bsky.social · 07/05/2025
Panel discussion on the current economic precarity of autonomous vehicle businesses. www.youtube.com/watch?v=gDG-... "We are at a really tough spot in generating flows of cash right now." 👇
youtube.com
The Future of AVs Panel | 2023 CCAT Symposium | Day 1
YouTube video by Center for Connected and Automated Transportation
111
Reposted by Theresa Eimer
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 14/04/2025
After a short era in which people questioned the value of academia in ML, its value is more obvious than ever. Big labs stopped publishing the minute commercial incentives showed up and are relentlessly focused on a singular vision of scaling. Academia is a meaningful complement, bringing... 1/2
221141
Reposted by Theresa Eimer
James MacGlashan @jmac-ai.bsky.social · 12/04/2025
It's strange to me that the focus of many people's worry is still "superintelligence" and not the reality we're currently living where increasingly authoritarian governments wield technology oppressively. This fantastical distraction based on speculative rhetoric is increasingly harmful.
0235
Reposted by Theresa Eimer
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 09/04/2025
A sensible perspective on humanoids in manufacturing (TLDR: if you can make humanoids, you can probably make better, more manufacturing specific things) blog.spec.tech/p/humanoid-r...
blog.spec.tech
Humanoid Robots in Manufacturing
Or, there's a reason we don't pull cars with mechanical horses
3588
Reposted by Theresa Eimer
EWRL @ewrl-org.bsky.social · 08/04/2025
Mark your calendars, EWRL is coming to Tübingen! 📅 When? September 17-19, 2025. More news to come soon, stay tuned!
03714
Reposted by Theresa Eimer
Nathan Lambert @natolambert.bsky.social · 07/04/2025
Llama 4 was a messy release: unreleased finetunes boosting scores, rumors of training on test, released on a weekend, etc As (open) models are commoditized / competition grows, what is the role of Meta's Llama efforts in the future? Should they continue?
buff.ly
Llama 4: Did Meta just push the panic button?
One of the weirdest releases of the year and understanding the future of the Llama endeavor. For the time being, we have some more amazing open weight models!
1379
Reposted by Theresa Eimer
Stephanie Brandl @stephaniebrandl.bsky.social · 07/04/2025
At least there is no need to jailbreak the model anymore 🫠 (Is there a counterpart to make it nicer 🎭?)
031
Theresa Eimer @theeimer.bsky.social · 03/04/2025
The school kids visiting me during this year's future day really had hard-hitting questions: "Do you still have a lot of free time?" Me, a pretty fresh and currently slightly overwhelmed PostDoc: "It's important to be good at time management. Like my colleague, maybe you should ask her."
020
Reposted by Theresa Eimer
Serge Belongie @serge.belongie.com · 03/04/2025
So far, 2,135 people have responded to the poll Søren and I posted a few days ago. Of those, 94.4% replied “Yes” to being interested in officially presenting accepted @neuripsconf.bsky.social papers in Europe. (1/7)
57924
Reposted by Theresa Eimer
Musa Okwonga @okwonga.bsky.social · 01/04/2025
German media I beg you one day just please go just one day without being obsessed with migration. One day. I promise it won’t kill you. You have lakes and mountains and good football and good healthcare and asparagus. You’ll be fine.
512414461
Theresa Eimer @theeimer.bsky.social · 01/04/2025
True, I've been "socialized" in the AutoML community, how to compare algorithms is a big deal there. I remember discussing with my advisor whether it's worth evaluating issues with improper HPO setup in RL, he thought it was so obvious that everyone must already be doing it (spoiler: not really)
010
Theresa Eimer @theeimer.bsky.social · 31/03/2025
Well, then there's only one alternative: "We define OurPO as PPO with lr=0.01, ent_coef=0.1.... and compare it to OurQN which is DQN with lr=...." 😂
010
Reposted by Theresa Eimer
Singularity's Bounty e/CC @catblanketflower.yuwakisa.com · 31/03/2025
Tell them their argument might be valid with different hyperparameters
1211
Theresa Eimer @theeimer.bsky.social · 31/03/2025
This obviously then also depends on budget, HPO method and combines performance and tunability into one score, but I think that's quite reasonable in practice. Not very satisfying for an empirical nihilist, though, I imagine 😉
210
Theresa Eimer @theeimer.bsky.social · 31/03/2025
Well, what validity are you looking for? The absolute "algorithm A is better than B on benchmark C" is hard wrt hyperparameters, but algorithm A is better than B on C given I can realistically try out 50 configurations" is what we often want in empirical ML anyway, no?
130
Reposted by Theresa Eimer
Gaël Varoquaux @gaelvaroquaux.bsky.social · 28/03/2025
So true, Gilles. Yes, it is a pretext task, but often, when we try real tasks, we find that the problems are not those we expected. We need more people looking at relevant problems. Kiri Wagstaff said this 15 years ago arxiv.org/abs/1206.4656
arxiv.org
Machine Learning that Matters
Much of current machine learning (ML) research has lost its connection to problems of import to the larger world of science and society. From this perspective, there exist glaring limitations in the d...
1204
Theresa Eimer @theeimer.bsky.social · 26/03/2025
My PhD supervisor discussed my first two or three reviews with me (including checking over wording etc.) and does that for all his PhDs, but I know that's not the standard in most other groups I'm familiar with...
010
Reposted by Theresa Eimer
Nick Erickson @nickerickson.bsky.social · 25/03/2025
We are excited to announce #FMSD: "1st Workshop on Foundation Models for Structured Data" has been accepted to #ICML 2025! Call for Papers: icml-structured-fm-workshop.github.io
01510
Reposted by Theresa Eimer
Pablo Samuel Castro @pcastr.bsky.social · 20/03/2025
I'm looking to hire a student researcher to work on an exciting project for 6 months in DeepMind Montreal. Requirements: - Full-time masters/PhD student 🧑🏾‍🎓 - Substantial expertise in multi-agent RL, ideally including publication(s) 🤖🤖 - Strong Python coding skills 🐍 Is this you? Get in touch!
33416
Reposted by Theresa Eimer
David Rügamer @davidruegamer.bsky.social · 19/03/2025
Still enough time to switch back to Vancouver for this year's @neuripsconf.bsky.social ? 😬
1132
Reposted by Theresa Eimer
Nathan Lambert @natolambert.bsky.social · 20/03/2025
We call for funding and support to open source players of all type, not just big tech companies.
0152
Reposted by Theresa Eimer
Stone Tao @stonet2000.bsky.social · 20/03/2025
some exciting news, ManiSkill/SAPIEN now has experimental support for MacOS for the CPU simulation and rendering. You can now do your local debugging/development on Mac. Example shown here is a Push-T policy trained on my 4090 running on my mac! try now: maniskill.readthedocs.io/en/latest/us...
0234
Theresa Eimer @theeimer.bsky.social · 19/03/2025
@amsks96.bsky.social @raghuspacerajan.bsky.social Didn't one of you look for something similar recently?
220
Theresa Eimer @theeimer.bsky.social · 19/03/2025
This is a great example where theory can explain an empirical effect and based on those theoretical results we get an actual fix as opposed to a band-aid. Super motivating when a lot of RL results, as valuable as they often are, still feel like guesswork!
140
Theresa Eimer @theeimer.bsky.social · 19/03/2025
Great blog (series) and great work! Both made me think about the panel at last year's ARLET workshop at ICML where I felt there was no really clear resolution of how to bring RL theory and empirical RL closer together.
191
Reposted by Theresa Eimer
Nathan Lambert @natolambert.bsky.social · 18/03/2025
A new harvard business school study showed that open-source software has crazy returns on investment. We have to work hard to build an open ecosystem for AI if we want the same to apply. Right now open models DO NOT have the same return as software, but they can. Study here: buff.ly/5xFjpvf
2113
Reposted by Theresa Eimer
Anne Urai @anne-urai.bsky.social · 17/03/2025
I think about this a lot. Thanks @behrenstimb.bsky.social for the wonderfully 90s-vibe blog full of wisdom! users.fmrib.ox.ac.uk/~behrens/Sta...
39213
Reposted by Theresa Eimer
Claire Vernade @claireve.bsky.social · 11/03/2025
I’ve put together a short list of opportunities for early career academics willing to come to Europe: www.cvernade.com/miscellaneou... This mostly covers France and Germany for now but I’m willing to extend it. I build on @ellis.eu resources and my own knowledge of these systems.
cvernade.com
Claire Vernade - European career opportunities
European Academic Career Opportunities in 2025
37526