Ben Edelman @benedelman.bsky.social · 20/05/2025Update: We are extending the MOSS workshop deadline to May 26th 4:59pm PDT (11:59pm UTC) 000
Ben Edelman @benedelman.bsky.social · 08/05/2025What if there were a workshop dedicated to *small-scale*, *reproducible* experiments? What if this were at ICML 2025? What if your submission (due May 22nd) could literally be a Jupyter notebook?? Pretty excited this is happening. Spread the word! sites.google.com/view/moss202... 262
Ben Edelman @benedelman.bsky.social · 17/01/20251/ Excited to share a new blog post from the U.S. AI Safety Institute! AI agents are becoming more capable, but they are vulnerable to prompt injections in external content – an agent may be given task A, but then be “hijacked” and perform malicious task B instead. www.nist.gov/news-events/...nist.govTechnical Blog: Strengthening AI Agent Hijacking EvaluationsLarge AI models are increasingly used to power agentic systems, or “agents,” which can automate complex tasks on behalf of users. 140
Ben Edelman @benedelman.bsky.social · 08/12/2024For years, this mysterious undulating loop has lived at the top of my personal homepage. 1152
Ben Edelman @benedelman.bsky.social · 02/12/2024I defended my PhD dissertation back in May. I didn't have time to share it widely then (newborn baby), but I think some of you might enjoy it, especially the opening chapters: benjaminedelman.com/assets/disse... 3313
Ben Edelman @benedelman.bsky.social · 29/11/20240/ I'd like to kick off my presence here with a question: why does learning work in practice? Why is the world such that we can we learn to predict things from other things in a computationally efficient way; why is "simplicity bias" empirically useful? Some explanations: 141