Sign in

vkrakovna.bsky.social

@vkrakovna.bsky.social
75 followers 26 following 7 posts

Research scientist in AI alignment at Google DeepMind. Co-founder of Future of Life Institute. Views are my own and do not represent GDM or FLI.

PostsRepliesMedia
vkrakovna.bsky.social @vkrakovna.bsky.social · 28/01/2026
Updates to the Google DeepMind AGI safety course: - New talk on Security & Control by @maryphuong.bsky.social - Updated talk on Interpretability by @neelnanda.bsky.social - Updated talk on Robust training & monitoring by @sebfar.bsky.social Check these out here: www.youtube.com/playlist?lis...
youtube.com
Google DeepMind AGI Safety Course - YouTube
A short course from Google DeepMind on AGI safety, covering alignment problems we can expect as AI capabilities advance, and our current approach to these pr...
020
vkrakovna.bsky.social @vkrakovna.bsky.social · 08/07/2025
As models advance, a key AI safety concern is deceptive alignment / "scheming" – where AI might covertly pursue unintended goals. Our paper "Evaluating Frontier Models for Stealth and Situational Awareness" assesses whether current models can scheme. arxiv.org/abs/2505.01420
161
Reposted by @vkrakovna.bsky.social
David Lindner @davidlindner.bsky.social · 04/04/2025
Super excited this giant paper outlining our technical approach to AGI safety and security is finally out! No time to read 145 pages? Check out the 10 page extended abstract at the beginning of the paper
163
vkrakovna.bsky.social @vkrakovna.bsky.social · 14/02/2025
We are excited to release a short course on AGI safety! The course offers a concise and accessible introduction to AI alignment problems and our technical / governance approaches, consisting of short recorded talks and exercises (75 minutes total). deepmindsafetyresearch.medium.com/1072adb7912c
deepmindsafetyresearch.medium.com
Introducing our short course on AGI safety
We are excited to release a short course on AGI safety for students, researchers and professionals interested in this topic. The course…
0185