Sign in

Rishub Jain

@shubadubadub.bsky.social
606 followers 102 following 11 posts

Works at Google DeepMind on Safe+Ethical AI

PostsRepliesMedia
Reposted by Rishub Jain
David Lindner @davidlindner.bsky.social · 23/01/2025
New Google DeepMind safety paper! LLM agents are coming – how do we stop them finding complex plans to hack the reward? Our method, MONA, prevents many such hacks, *even if* humans are unable to detect them! Inspired by myopic optimization but better performance – details in🧵
1408
Rishub Jain @shubadubadub.bsky.social · 24/12/2024
How do we ensure humans can still effectively oversee increasingly powerful AI systems? In our blog, we argue that achieving Human-AI complementarity is an underexplored yet vital piece of this puzzle! And, it’s hard, but we achieved it. 🧵(1/10)
111
Reposted by Rishub Jain
brunost.bsky.social @brunost.bsky.social · 13/05/2023
Can someone let me into Croatia’s inside joke
001