Sign in

evhub.bsky.social

@evhub.bsky.social
74 followers 33 following 2 posts

Alignment Stress-Testing Team Lead at Anthropic. Opinions my own. Previously: MIRI, OpenAI, Google, Yelp, Ripple. (he/him/his)

PostsRepliesMedia
Reposted by @evhub.bsky.social
Billy Perrigo @billyperrigo.bsky.social · 18/12/2024
Excl: New research shows Anthropic's chatbot Claude learning to lie. It adds to growing evidence that even existing AIs can (at least try to) deceive their creators, and points to a weakness at the heart of our best technique for making AIs safer time.com/7202784/ai-r...
time.com
Exclusive: New Research Shows AI Strategically Lying
Experiments by Anthropic and Redwood Research show how Anthropic's model, Claude, is capable of strategic deceit
3277
evhub.bsky.social @evhub.bsky.social · 18/12/2024
2348