Karim Abdel Sadek @karimabdel.bsky.social · 08/07/2025*New Paper* 🚨 Goal misgeneralization occurs when AI agents learn the wrong reward function, instead of the human's intended goal. 😇 We show that training with a minimax regret objective provably mitigates it, promoting safer and better-aligned RL policies! 192
Reposted by Karim Abdel SadekJoel Z Leibo @jzleibo.bsky.social · 21/02/2025CAIF's new and massive report on multi-agent AI risks will be really useful resource for the field www.cooperativeai.com/post/new-rep...cooperativeai.comCooperative AIPlaintext Code Block 031
Reposted by Karim Abdel SadekEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 18/02/2025A large group of us (spearheaded by Denizalp Goktas) have put out a position paper on paths towards foundation models for strategic decision-making. Language models still lack these capabilities so we'll need to build them: hal.science/hal-04925309... 2337
Reposted by Karim Abdel SadekEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 14/02/2025Model-free deep RL algorithms like NFSP, PSRO, ESCHER, & R-NaD are tailor-made for games with hidden information (e.g. poker). We performed the largest-ever comparison of these algorithms. We find that they do not outperform generic policy gradient methods, such as PPO. arxiv.org/abs/2502.08938 1/N 39321
Reposted by Karim Abdel SadekVincent Conitzer @conitzer.bsky.social · 09/01/2025The 2025 Cooperative AI summer school (9-13 July 2025 near London) is now accepting applications, due March 7th! www.cooperativeai.com/summer-schoo...cooperativeai.comCooperative AI 1155
Reposted by Karim Abdel SadekEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 21/12/2024The magic thing that humans do is a pretty good job at solving tasks under high uncertainty about the problem specification. We also frequently are capable of doing this collaboratively. I still do not see evidence that models can do any part of this. 68211
Karim Abdel Sadek @karimabdel.bsky.social · 08/12/2024I will be at @neuripsconf.bsky.social this week! Would love to chat about Multi-agent systems, RL, Human-AI Alignment, or anything interesting :) I'm also applying for PhD programs this cycle, feel free to reach out for any advice! More about me: karim-abdel.github.io 092
Reposted by Karim Abdel SadekClément Canonne @ccanonne.github.io · 18/11/2024I give you a loaded coin, with some (unknown) probability 0<p<1 of landing Heads, and I ask you to generate a fair coin toss. Great! We know how to do this! This is the Von Neumann trick: toss twice. If HH or TT, repeat; if HT or TH, return the first. Problem solved? Not quite... This can be bad! 3417