Sign in

Ben Edelman

@benedelman.bsky.social
209 followers 36 following 32 posts

Thinking about how/why AI works/doesn't, and how to make it go well for us. Currently: AI Agent Security @ US AI Safety Institute benjaminedelman.com

PostsRepliesMedia
Ben Edelman @benedelman.bsky.social · 20/05/2025
Update: We are extending the MOSS workshop deadline to May 26th 4:59pm PDT (11:59pm UTC)
000
Ben Edelman @benedelman.bsky.social · 08/05/2025
What if there were a workshop dedicated to *small-scale*, *reproducible* experiments? What if this were at ICML 2025? What if your submission (due May 22nd) could literally be a Jupyter notebook?? Pretty excited this is happening. Spread the word! sites.google.com/view/moss202...
262
Ben Edelman @benedelman.bsky.social · 17/01/2025
1/ Excited to share a new blog post from the U.S. AI Safety Institute! AI agents are becoming more capable, but they are vulnerable to prompt injections in external content – an agent may be given task A, but then be “hijacked” and perform malicious task B instead. www.nist.gov/news-events/...
nist.gov
Technical Blog: Strengthening AI Agent Hijacking Evaluations
Large AI models are increasingly used to power agentic systems, or “agents,” which can automate complex tasks on behalf of users.
140
Ben Edelman @benedelman.bsky.social · 08/12/2024
For years, this mysterious undulating loop has lived at the top of my personal homepage.
1152
Ben Edelman @benedelman.bsky.social · 02/12/2024
I defended my PhD dissertation back in May. I didn't have time to share it widely then (newborn baby), but I think some of you might enjoy it, especially the opening chapters: benjaminedelman.com/assets/disse...
3313
Ben Edelman @benedelman.bsky.social · 29/11/2024
0/ I'd like to kick off my presence here with a question: why does learning work in practice? Why is the world such that we can we learn to predict things from other things in a computationally efficient way; why is "simplicity bias" empirically useful? Some explanations:
141