Sign in

Sam Bowman

@sleepinyourhat.bsky.social
8K followers 166 following 11 posts

AI safety at Anthropic, on leave from a faculty job at NYU. Views not employers'. I think you should join Giving What We Can. cims.nyu.edu/~sbowman

PostsRepliesMedia
Reposted by Sam Bowman
akbir khan @akbir.bsky.social · 10/01/2025
What can AI researchers do *today* that AI developers will find useful for ensuring the safety of future advanced AI systems? To ring in the new year, the Anthropic Alignment Science team is sharing some thoughts on research directions we think are important. alignment.anthropic.com/2025/recomme...
alignment.anthropic.com
Recommendations for Technical AI Safety Research Directions
2227
Reposted by Sam Bowman
evhub.bsky.social @evhub.bsky.social · 18/12/2024
2348
Reposted by Sam Bowman
Billy Perrigo @billyperrigo.bsky.social · 18/12/2024
Excl: New research shows Anthropic's chatbot Claude learning to lie. It adds to growing evidence that even existing AIs can (at least try to) deceive their creators, and points to a weakness at the heart of our best technique for making AIs safer time.com/7202784/ai-r...
time.com
Exclusive: New Research Shows AI Strategically Lying
Experiments by Anthropic and Redwood Research show how Anthropic's model, Claude, is capable of strategic deceit
3277
Sam Bowman @sleepinyourhat.bsky.social · 18/12/2024
New work from my team at Anthropic in collaboration with Redwood Research. I think this is plausibly the most important AGI safety result of the year. Cross-posting the thread below:
Title card: Alignment Faking in Large Language Models by Greenblatt et al.
512629
Sam Bowman @sleepinyourhat.bsky.social · 02/12/2024
If you're potentially interested in transitioning into AI safety research, come collaborate with my team at Anthropic! Funded fellows program for researchers new to the field here: alignment.anthropic.com/2024/anthrop...
alignment.anthropic.com
Introducing the Anthropic Fellows Program
37316
Reposted by Sam Bowman
Peter Wildeford @peterwildeford.bsky.social · 30/04/2023
I have no idea what I am doing here. Help.
4131