Sign in

Willem Röpke

@willemropke.bsky.social
1.4K followers 414 following 59 posts

PhD student | Interested in all things decision-making and learning

PostsRepliesMedia
Willem Röpke @willemropke.bsky.social · 13/05/2025
010
Willem Röpke @willemropke.bsky.social · 22/04/2025
I think the Qwen team is missing up on a huge opportunity to basically be the default model in all neurips submissions by not releasing Qwen3
120
Willem Röpke @willemropke.bsky.social · 10/04/2025
Using LLMs to come up with prompts for LLMs to then ask the LLMs to then train the LLMs to then ....
020
Willem Röpke @willemropke.bsky.social · 09/04/2025
Manifesting Qwen 3
010
Willem Röpke @willemropke.bsky.social · 04/04/2025
RIP to my investments from the past few years, it was nice seeing the green while it lasted
000
Willem Röpke @willemropke.bsky.social · 02/04/2025
The people demand Qwen3!
000
Willem Röpke @willemropke.bsky.social · 24/03/2025
I've been bashing my head against a wall trying to make TRL and their new vllm-serve work and holy moly it's just an infinite pain why must i suffer
000
Willem Röpke @willemropke.bsky.social · 22/03/2025
Why does reading a book feel so much more satisfying than watching a TV show? Both are ways of consuming content so I don't get the difference
000
Willem Röpke @willemropke.bsky.social · 12/03/2025
Bought a cherry coke on accident today. Horrible things happening everywhere apparently
010
Willem Röpke @willemropke.bsky.social · 12/03/2025
This is actually insanely clever, I would've never thought about this. Seems very interesting and important to fix!
000
Willem Röpke @willemropke.bsky.social · 28/02/2025
I don't recall seeing a video in the recent past that depressed me as much as what I just watched unfolding in the Oval Office
020
Willem Röpke @willemropke.bsky.social · 17/02/2025
Exciting news! My paper on multi-objective reinforcement learning was accepted at AAMAS 2025! We introduce IPRO (Iterated Pareto Referent Optimisation)—a principled approach to solving multi-objective problems. 🔗 Paper: arxiv.org/abs/2402.07182 💻 Code: github.com/wilrop/ipro
2265
Willem Röpke @willemropke.bsky.social · 12/02/2025
This is unholy
030
Willem Röpke @willemropke.bsky.social · 12/02/2025
How can I stop ChatGPT from talking to me with emojis, this is just the worst update I've ever experienced. I've put it in its memory, in my details, and I even repeat it in the chat but it's just replying like 👉🥺👈
000
Willem Röpke @willemropke.bsky.social · 11/02/2025
Macron is the goat French people don't appreciate true genius
110
Willem Röpke @willemropke.bsky.social · 11/02/2025
Why did OpenAI update chatGPT to use emojis in its responses? I hate it and even when I explicitly say this it just keeps doing it.
000
Willem Röpke @willemropke.bsky.social · 05/02/2025
To whomever put my email in some spam list: I fart in your general direction
000
Willem Röpke @willemropke.bsky.social · 04/02/2025
The fact that in the year 2025 we are still dealing with the stupid "make the paper fit in an arbitrary format for the camera ready submission" minigame is killing me. Either let me group authors or let me put acknowledgements after the main text. This isn't hard.
240
Willem Röpke @willemropke.bsky.social · 31/01/2025
Does anyone have any good hacks for making the AAMAS template not suck for people with multiple affiliations? I lose a gazillion lines for basically no reason...
100
Willem Röpke @willemropke.bsky.social · 29/01/2025
I found a very promising open problem in AI Computing a MEDIAN over a list of rows where one of the elements is just an empty array
010
Willem Röpke @willemropke.bsky.social · 27/01/2025
I think this is the best paper I’ve ever read: arxiv.org/abs/2404.03715 A strong emphasis on theoretically principled algorithms for RLHF followed by motivated practical implementations. Well-written and a clear overview of the relevant background and related work. 10/10 no comments
arxiv.org
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
This paper studies post-training large language models (LLMs) using preference feedback from a powerful oracle to help a model iteratively improve over itself. The typical approach for post-training L...
040
Willem Röpke @willemropke.bsky.social · 20/01/2025
Deepseek making my day just a little better
020
Willem Röpke @willemropke.bsky.social · 20/01/2025
I realise I'm woefully unqualified on this topic, but can someone please explain why we still don't have personal carrier drones? This seems like an obvious next step in transportation and given the state of our tech tree shouldn't be that hard?
110
Willem Röpke @willemropke.bsky.social · 15/01/2025
I think we should do congestion pricing in a lot more places
050
Willem Röpke @willemropke.bsky.social · 13/01/2025
Claude just declined my attempt at bribing it to do a better job. Not sure whether to be happy or sad
110
Willem Röpke @willemropke.bsky.social · 08/01/2025
I learned to stop reading documentation and just ask ChatGPT So far seems to work out great
100
Willem Röpke @willemropke.bsky.social · 03/01/2025
I just cooked a chatgpt recipe from some leftovers in my fridge and I gotta say it was delicious. The future is now
110
Willem Röpke @willemropke.bsky.social · 02/01/2025
Can someone please convince me that buying a 3D printer while living in a small appartement is a good idea?
430
Willem Röpke @willemropke.bsky.social · 25/12/2024
I'm having a weird problem with training DQN on minatar (specifically the gymnax version). In space invaders and breakout, my eval metrics are extremely unstable while my train metric is very smooth. See an example of space invaders below (eval left, train right). Any ideas of what went wrong?
101
Willem Röpke @willemropke.bsky.social · 17/12/2024
I just learned that this is allowed in Python. Who do I talk to to get this banned?
200
Willem Röpke @willemropke.bsky.social · 16/12/2024
I just spent 1h+ trying to solve an annoying issue which came down to downgrading numpy+tensorflow+keras feels great
030
Willem Röpke @willemropke.bsky.social · 12/12/2024
I just made a commit that fixed a typo with the message "fi typo" 🤦‍♂️
020
Willem Röpke @willemropke.bsky.social · 11/12/2024
Is there a rule of thumb for RL algorithms that use a replay buffer for determining the size of this buffer relative to the total number of timesteps? For example if DQN takes 500k steps, the RB should be of size ... It could also depend on other parameters, just looking for a general rule of thumb.
310
Willem Röpke @willemropke.bsky.social · 05/12/2024
If any of the cool industry labs want to open an RL (or any ML topic tbh) lab in Brussels in the next year or so, I'd greatly appreciate it! I know someone (me) that wants to continue research but is quite keen on sticking around in Belgium...
150
Willem Röpke @willemropke.bsky.social · 02/12/2024
Back from my vacation! Did I miss any cool papers or other work? Also, Berlin is really amazing!
020
Willem Röpke @willemropke.bsky.social · 22/11/2024
Okay, since a lot of RL people have migrated over here I'm going to do a small experiment! Please drop your favorite RLHF or preference-based RL papers here. I want to speedrun a lit review for my next project!
4202
Willem Röpke @willemropke.bsky.social · 21/11/2024
Launching a sweep on wandb and seeing 15 runs 1 minute later is true nightmare fuel Every project I start is so much fun until it's time to run experiments...
200
Willem Röpke @willemropke.bsky.social · 21/11/2024
Is there a consensus on the best way to use attention layers in RL? In particular, I want to somehow use it as part of my encoder that will later feed into other components (e.g. the policy, critic, whatever)
000
Willem Röpke @willemropke.bsky.social · 21/11/2024
If anyone wants to put me on a starter pack I'm: - super funny - really handsome - do a bit of RL on the side
140
Willem Röpke @willemropke.bsky.social · 20/11/2024
My favorite bug is the one you just solved but forgot to pull on the cluster where you are actually running your experiments so much fun not at all the worst ever
010
Willem Röpke @willemropke.bsky.social · 19/11/2024
Follow me for amazing content about machine learning and reinforcement learning. (testing to see if I can get more followers on the new place than on twitter)
160