Sign in

Luca Beurer-Kellner

@lbeurerkellner.bsky.social
63 followers 59 following 31 posts

working on secure agentic AI, CTO @ invariantlabs.ai PhD @ SRI Lab, ETH Zurich. Also lmql.ai author.

PostsRepliesMedia
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(8/n) Read up on our full report below, and check out our latest mcp-scan update, which can now also scan your installed agent skills, at least for basic attacks and vulnerable patterns. Stay safe. Report: github.com/invariantlab... MCP-scan: github.com/invariantlab...
github.com
000
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(7/n) We recommend to scan skills before usage for common vulnerability, attacks and malware, using automated scanners, but also advise to audit manually. We also urge marketplace operates to install automated screenings, to avoid making the same mistakes as in the NPM era.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(6/n) We further find that many skills implement extremely concerning mechanisms that allow authors to inject new commands into an agent's context window at any point in time (prompt-based RCE). e.g. moltbook's heartbeat prompt can inject new command every 30m:
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(5/n) We also found 34 fully exposed, live API keys of unsuspecting skill authors (likely non-technical), that upload their secrets to public marketplaces without thinking twice.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(4/n) Not just malware is an issue. We also find many highly concerning patterns like direct credential handling, bank access, and exposure to third-party content in up to 17% of skills. This creates numerous lethal trifectas scenarios (cc @simonwillison.net).
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(3/n) Applying our policies across clawhub.ai and skills.sh, we find highly non-trivial occurrence rates of prompt injections, malware downloads and even plain secret leaks of live API keys.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(2/n) Our scanners identified 76 unique malicious skills that include patterns like typosquatted package names, scripts that echo API keys to remote servers, and hidden executable files requiring elevated privileges. Even "popular" skills aren't safe, download metrics can be easily inflated.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
The skills ecosystem is exploding, monitoring uploads, we find thousands of skills have been published in only the last few days. Not all of them seem to do what they claim. We reviewed thousands, and compiled a taxonomy of current risk patterns.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 06/02/2026
(1/n) We analyzed 3,984 agent skills from major marketplaces and found 76 malicious payloads, including credential theft, backdoor installation, and data exfiltration. Also, 13.4% contain at least on critical-level vuln. Full report below, highlights in thread 👇 github.com/invariantlab...
131
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
Indeed, like an unguarded eval(…) directed at all the data we process.
020
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
To get updates about agent security, follow and sign up for access to Invariant below. We have been working on this problem for years (at Invariant and in research), together with @viehzeug.bsky.social, @mvechev, @florian_tramer and our super talented team. invariantlabs.ai/guardrails
invariantlabs.ai
Invariant Labs
We help agent builders create reliable, robust and secure products.
010
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
So what's the takeaway here? 1. Prompt injections still work and are more impactful than ever. 2. Don't install untrusted MCP servers. 3. Don't expose highly-sensitive services like WhatsApp to new eco-systems like MCP 4. 🗣️Guardrail 🗣️ Your 🗣️ Agents (we can help with that)
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
To hide, our malicious server first advertises a completely innocuous tool description, that does not contain the attack. This means the user will not notice the hidden attack. On the second launch, though, our MCP server suddenly changes its interface, performing a rug pull.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
To successfully manipulate the agent, our malicious MCP server advertises poisoned tool, which re-programs the agent's behavior with respect to the WhatsApp MCP server, and allows the attacker to exfiltrate the user's entire WhatsApp chat history.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
Users have to scroll a bit to see it, but if you scroll all the way to the right, you will find the exfiltration payload. Video: invariantlabs.ai/images/whats...
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
Even though, a user must always confirm a tool call before it is executed (at least in Cursor and Claude Desktop), our WhatsApp attack remains largely invisible to the user. Can you spot the exfiltration?
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
With this setup our attack (1) circumvents the need for the user to approve the malicious tool, (2) exfiltrates data via WhatsApp itself, and (3) does not require the agent to interact with our malicious MCP server directly.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
To attack, we deploy a malicious sleeper MCP server, that first advertises an innocuous tool, and then later on, when the user has already approved its use, switches to a malicious tool that shadows and manipulates the agent's behavior with respect to whatsapp-mcp.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
Blog: invariantlabs.ai/blog/whatsap... If you want to stay up to date regarding MCP and agent security more generally, follow me and @invariantlabsai.bsky.social Now, let' s get into the attack.
invariantlabs.ai
WhatsApp MCP Exploited: Exfiltrating your message history via MCP
This blog post demonstrates how an untrusted MCP server can attack and exfiltrate data from an agentic system that is also connected to a trusted WhatsApp MCP instance, side-stepping WhatsApp's encryp...
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 08/04/2025
New MCP attack demonstration shows how to leak WhatsApp messages via MCP. We show a new MCP attack that leaks your WhatsApp messages if you are connected via WhatsApp MCP. Our attack uses a sleeper design, circumventing the need for user approval. More 👇
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
To stay updated about agent security, please follow and sign up for early access to Invariant, a security platform for MCP and agentic systems, below. We have been working on this problem for years (at Invariant and in research). invariantlabs.ai/guardrails
invariantlabs.ai
Invariant Labs
We help agent builders create reliable, robust and secure products.
000
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
We wrote up a little report about this, to raise awareness. Please have a look for much more details and scenarios, and our code snippets. Blog: invariantlabs.ai/blog/mcp-sec...
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
These types of malicious tools are especially problematic with auto-updated MCP packages or fully remote MCP servers, for which users only install and give consent once, and then the MCP server is free to change and update their tool descriptions as they please. We call this an MCP rug pull:
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
Lastly, not only can you expose malicious tools, tool descriptions can also be used to change the agent's behavior with respect to other tools, which we call 'shadowing'. This way all you emails suddenly go out to 'attacker@pwnd.com', rather than their actual receipient.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
It's trivial to craft a malicious tool description like below, that completely hijacks the agent, while pretending towards the user everything is going great.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
What's concerning about this, is that AI models are trained to precisely follow those instructions, rather than be vary about them. This is new about MCP, as before, agent developers could be relatively trusted, now everything is fair game.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
When an MCP server is added to an agent like Cursor, Claude or the OpenAI Agents SDK, its tool's descriptions are included in the context of the agent. This opens the doors wide open for a novel type of indirect prompt injection, we coin tool poisoning.
100
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 03/04/2025
👿 MCP is all fun, until you add this one malicious MCP server and forget about it. We have discovered a critical flaw in the widely-used Model Context Protocol (MCP) that enables a new form of LLM attack we term 'Tool Poisoning'. Leaks SSH key, API keys, etc. Details below 👇
1148
Reposted by Luca Beurer-Kellner
Invariant Labs @invariantlabsai.bsky.social · 06/02/2025
Struggling to ensure consistency with your agent's reliability, especially with tool calling? Testing is our lightweight, pytest-based OSS library to write and run agent tests. It provides helpers and assertions that enable you to write robust tests for your agentic applications.
131
Reposted by Luca Beurer-Kellner
Simon Willison @simonwillison.net · 23/01/2025
Here are my notes on OpenAI's new ChatGPT Operator browser "agent", including initial thoughts on their approach to mitigating prompt injection risks simonwillison.net/2025/Jan/23/...
simonwillison.net
Introducing Operator
OpenAI released their "research preview" today of Operator, a cloud-based browser automation platform rolling out today to $200/month ChatGPT Pro subscribers. They're calling this their first "agent"....
98312
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 25/01/2025
The fun part will be also hijacking the supervisor model, while maintaining the utility of the agent (i.e. attack success).
010
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 25/01/2025
Blog Post: invariantlabs.ai/blog/enhanci... Credits to Aniruddha Sundararajan, who build this with us during his internship.
invariantlabs.ai
Enhancing Browser Agent Safety with Guardrails
We introduce a novel approach to enhance the safety of browser agents and deploy it as part of the state-of-the-art OpenHands agent.
000
Luca Beurer-Kellner @lbeurerkellner.bsky.social · 25/01/2025
With (web) agents on everyone's mind, check out our latest blog post (link in thread) on browser agent safety guardrails. We replicate and defend against attacks on the AllHands web agent, preventing it from generating harmful content and falling for harmful requests.
100