Sign in

Simon Lermen

@simonlermen.bsky.social
159 followers 332 following 33 posts

I work on AI safety and AI in cybersecurity

PostsRepliesMedia
Simon Lermen @simonlermen.bsky.social · 23/07/2026
I helped run this human study (n=4,100) on how AI can automate voice phishing. We found that humans can't tell the best AI voice models apart from humans anymore. substack.com/home/post/p-...
substack.com
AI Voice Phishing Performs on Par With Human Scammers at a Fraction of the Cost
TL;DR: We ran a large-scale human-subject study (n=4,100) to measure susceptibility to AI-powered voice phishing, using six leading AI voice models.
000
Simon Lermen @simonlermen.bsky.social · 13/05/2026
Some are claiming through Sudan xcancel.com/TyskNIP/stat...
001
Simon Lermen @simonlermen.bsky.social · 18/03/2026
Interesting article by the nytimes featuring our research on anonymity.
000
Simon Lermen @simonlermen.bsky.social · 18/03/2026
Thanks for featuring our research.
000
Simon Lermen @simonlermen.bsky.social · 12/03/2026
What's the endstate of AI making mass surveillance cheap and scalable? "The study [..] speaks to one important aspect of the shifting paradigm of mass surveillance. (The research was led by Simon Lermen of MATS Research and Daniel Paleka from ETH Zurich.)" www.bloomberg.com/opinion/arti...
bloomberg.com
Anthropic Isn’t Exaggerating About an AI Panopticon
In the debate about the military’s use of artificial intelligence, prompted by Anthropic’s dispute with the Pentagon that’s now headed to the courts, much has been said about the concerns related to a...
000
Simon Lermen @simonlermen.bsky.social · 05/03/2026
What a video. Crazy that it's like one 84 year old politician who is taking this stuff seriously.
010
Reposted by Simon Lermen
koenfucius @koenfucius.bsky.social · 25/02/2026
“I didn’t write that” “Yes you did” Research by @simonlermen.bsky.social et al shows LLMs can deanonymize pseudonymous users of online platforms using unstructured content (eg link pseudonymous Hacker News posts with LinkedIn profiles or interview transcripts): buff.ly/bAdgQpx
011
Simon Lermen @simonlermen.bsky.social · 25/02/2026
I am one of the authors. Also check out my blogpost: simonlermen.substack.com/p/large-scal...
simonlermen.substack.com
Large-Scale Online Deanonymization with LLMs
We measure the capabilities of LLMs to deanonymize users online.
040
Simon Lermen @simonlermen.bsky.social · 20/02/2026
Happy to share my matsprogram.org project that I have been working on in the last couple of months. We explore how LLMs can be used for large-scale deanonymization online.
040
Simon Lermen @simonlermen.bsky.social · 04/07/2025
Our paper on AI-powered spear phishing, co-authored with @fredheiding.bsky.social , has been accepted at the ICML 2025 Workshop on Reliable and Responsible Foundation Models! openreview.net/pdf?id=f0uFp...
openreview.net
011
Simon Lermen @simonlermen.bsky.social · 20/04/2025
Do you think there is any comparable thing in China to AI Twitter or Bluesky? Where people discuss ideas
100
Simon Lermen @simonlermen.bsky.social · 20/04/2025
Are you working at DeepSeek?
000
Simon Lermen @simonlermen.bsky.social · 28/02/2025
Why so mean old man
100
Simon Lermen @simonlermen.bsky.social · 25/02/2025
Grok's DeepSearch was launched with Zero safety features, you can ask it about assasslnations, dru*gs. This has been online for a few days now with no changes.
020
Simon Lermen @simonlermen.bsky.social · 23/01/2025
I’m mostly interested in not dying
020
Simon Lermen @simonlermen.bsky.social · 22/01/2025
If you are trying to understand its reasoning, it seems like a necessary step to have legible chain-of-thought.
120
Simon Lermen @simonlermen.bsky.social · 15/01/2025
you should be carefully here, huge datacenters with their own powerstructures are being discussed, huge new semiconductor facilities. situation might change openai.com/global-affai...
openai.com
OpenAI’s Economic Blueprint
The Blueprint outlines policy proposals for how the US can maximize AI’s benefits, bolster national security, and drive economic growth
100
Simon Lermen @simonlermen.bsky.social · 15/01/2025
To be fair, the pre-training and all those mega datacenters do have some significant environmental impact. buying products from AI labs does fund this. But agree that individual energy use per reply is like the weakest argument against AI.
110
Simon Lermen @simonlermen.bsky.social · 04/01/2025
I published a human study with @fredheiding.bsky.social We use AI agents built from GPT-4o and Claude 3.5 Sonnet to search the web for available information on a target and use this for highly personalized phishing messages. achieved click-through rates above 50% www.lesswrong.com/posts/GCHyDK...
lesswrong.com
Human study on AI spear phishing campaigns — LessWrong
TL;DR: We ran a human subject study on whether language models can successfully spear-phish people. We use AI agents built from GPT-4o and Claude 3.5…
041
Simon Lermen @simonlermen.bsky.social · 02/01/2025
Has anyone ever tried with constitutional AI to add something on: always show your entire reasoning? What happens if you ask the model if it left out steps in its reasoning? can it verbalize them?
100
Simon Lermen @simonlermen.bsky.social · 28/12/2024
They achieve this in part by immediately releasing models after training such as o3, other companies wait for safety and security evaluations and estimates of societal impact. They also used to wait with releases such as with GPT-4
110
Simon Lermen @simonlermen.bsky.social · 27/12/2024
sometimes fancy terms just serve to confuse people
030
Simon Lermen @simonlermen.bsky.social · 27/12/2024
They have already made billions in revenue, but defining it as profits makes it almost impossible to reach
001
Simon Lermen @simonlermen.bsky.social · 27/12/2024
crazy that they use profits instead of revenue. so they can always just hack this by spending a bit more on R&D
100
Simon Lermen @simonlermen.bsky.social · 27/12/2024
my guess is he thinks of some sort of conscious experience of wanting here...
010
Simon Lermen @simonlermen.bsky.social · 26/12/2024
its behavior is at if it wants to win, same will be true about powerful AI agents. whether it actually wants something in a way that satisfies you doesn't matter
000
Simon Lermen @simonlermen.bsky.social · 26/12/2024
So RL-training the model to achieve some goal such as with constitutional AI can't lead to the model having a goal? do you think AlphaZero wants to win at chess?
100
Simon Lermen @simonlermen.bsky.social · 22/12/2024
💯
010
Simon Lermen @simonlermen.bsky.social · 16/12/2024
Well, we observe computation in superposition
100
Simon Lermen @simonlermen.bsky.social · 16/12/2024
I agree that it doesn't PROVE multiverses. But I don't like the sneering tone, what is superposition? It sure seems like the electron is in many places at once, all interpretations of that seem a bit crazy. Everett's manyworlds is a common position among physicists, including some i know.
200
Simon Lermen @simonlermen.bsky.social · 16/12/2024
The many worlds interpretation is a commonly held view by many physicists. And it is not like other interpretations are less "weird".
000
Simon Lermen @simonlermen.bsky.social · 16/12/2024
The many worlds interpretation is a commonly held view by many physicists. And it is not like other interpretations are less "weird"
200
Simon Lermen @simonlermen.bsky.social · 14/12/2024
I don't understand why we don't have more conferences in countries with easy visa policies
110
Simon Lermen @simonlermen.bsky.social · 13/12/2024
I'll be at the SafeGenAI workshop on Sunday presenting on research I did on safety in AI agents. I will talk about results from these two blog posts: www.lesswrong.com/posts/ZoFxTq... And: www.lesswrong.com/posts/Lgq2Dc...
lesswrong.com
Current safety training techniques do not fully transfer to the agent setting — LessWrong
TL;DR: We are presenting three recent papers which all share a similar finding, i.e. the safety training techniques for chat models don’t transfer we…
040
Reposted by Simon Lermen
Arthur Conmy @arthurconmy.bsky.social · 22/11/2024
I'm very bullish on automated research engineering soon, but even I was surprised that AI agents are twice as good as humans with 5+ years of experience or from a top AGI or safety lab at doing tasks in 2 hours. Paper: metr.org/AI_R_D_Evalu...
metr.org
181