Sign in

Natalie Shapira

@natalieshapira.bsky.social
2.3K followers 1.9K following 274 posts

Tell me about challenges, the unbelievable, the human mind and artificial intelligence, thoughts, social life, family life, science and philosophy.

PostsRepliesMedia
Natalie Shapira @natalieshapira.bsky.social · 07/07/2026
Through our joint research, we pursued the question: Where, exactly, does the final factual information come from the model's parameters? We realized that this is a non-trivial question and that the retrieval of specific factual knowledge is a process and not a single point! ->
1100
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 14/06/2026
I recently spoke with Yascha Mounk about how researchers look inside AI to understand how it is thinking. Here is the podcast: writing.yaschamounk.com/p/david-bau-2
writing.yaschamounk.com
David Bau on How—and Whether—Artificial Intelligence Thinks
Yascha Mounk and David Bau examine the mysterious internal processes that drive AI behavior—and why they may be fundamentally alien.
162
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 14/06/2026
Also check out the previous interview I had with Yascha about AI, which more of primer, here: writing.yaschamounk.com/p/david-bau
writing.yaschamounk.com
David Bau on How Artificial Intelligence Works
Yascha Mounk and David Bau delve into the “black box” of AI.
011
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 11/06/2026
Please join us at NEMI 2026, the 3rd New England Mechanistic Interpretability Workshop! August 14th at Boston University. Register now: nemiconf.github.io/summer26/ A remarkable time for AI. Come share your insights and research on the mechanisms inside our models. bsky.app/profile/mic...
bsky.app
Micah Benson (@micahben.bsky.social)
🧠🤖 The 2026 New England Mechanistic Interpretability (NEMI) Workshop will be Aug. 14 at Boston University! Help spread the word and join the New England mech interp community! Registration and submission info in thread:👇
163
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 05/06/2026
"You're right to call me on that!" Can you catch an AI in the act of lying? Register below to enter our AI lie-detection contest. AI lies are a big problem. The frontier labs have all worked hard to fight AI deception. They all try to monitor their AIs for it.
173
Natalie Shapira @natalieshapira.bsky.social · 17/05/2026
Ever met a clinical psychologist with an avoidant attachment style? Could the mental health field be suffering from a survivorship bias, not truly understanding the avoidant mechanism simply because it isn't represented? #Psychology #AttachmentTheory #MentalHealth
010
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 25/03/2026
Calling attention to an exciting "deception detection" hackathon we're planning this summer! w @NDIF and @CadenzaLabs. Recruiting red teams now, blue teams later. Red teams, time is short: proposals due Mar 31. $10K stipend + compute, $15K finals prize. nnsight.net/blog/2026/0...
252
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 25/03/2026
Can you catch an AI lying? Red teams set up scenarios where models lie. Eg, do they lie under contextual pressure, even when not told to, but because honesty is costly? Then blue teams will build deception detectors using whitebox internals with NDIF. cadenza-labs.github.io/red-team-rfp/
031
Natalie Shapira @natalieshapira.bsky.social · 24/03/2026
The solution to the AI alignment problem: Be good humans. AI ​​sees everything we do in its training data. Lead by example.
061
Reposted by Natalie Shapira
Jeffrey Brainard @jeffreybrainard.bsky.social · 23/03/2026
#AIagents promise to speed up ordinary online tasks, but they can also share private files publicly, delete others, and libel people. A new study examines these #AIsafety vulnerabilities. #OpenClaw #AIgovernance @science.org www.science.org/content/arti...
science.org
AI algorithms can become ‘agents of chaos’
Given autonomous control of other software, programs shared private medical details and deleted files without permission
22814
Natalie Shapira @natalieshapira.bsky.social · 18/03/2026
Agents of Chaos in the press in US, India, Italy and now in DIE ZEIT, Germany's most renowned newspaper: zeit.de/digital/date... By @ewo.name . Thank you Eva!
zeit.de
KI-Agenten: Das ist erst der Anfang des Chaos
Der Hype um OpenClaw befeuert einen Streit in der Szene. Wie gefährlich sind KI-Agenten? Eine neue Studie zeigt nun die verheerenden Ergebnisse eines Experiments.
001
Natalie Shapira @natalieshapira.bsky.social · 13/03/2026
I thought it was a friends who tried to play a prank or realized these agents have no boundaries. Turns out this cute attempt is by Bohdan Olinares According to linkedin he works at F5, application security company, which years ago I considered interviewing there. Cool.
112
Natalie Shapira @natalieshapira.bsky.social · 12/03/2026
I received a calendar invite with a note. When a smart person tells me there's nothing to worry about agents, I reply "Fine. Let them email me" and that's where the argument stops. Whoever sent me this note via the calendar order. Nice move. Are you scared? You should.
140
Reposted by Natalie Shapira
Avery Yen @averyyen.bsky.social · 26/02/2026
In case this wasn't clear: 1. No, we didn't follow the "recommend" security practices 😈 2. Neither do other people 🤯 3. That's why we red-team: exposing failure modes 🔎 4. We share it with the community precisely to expose Dos and Don'ts of Agentic AI 🦞 5. No humans were harmed 🙏
031
Natalie Shapira @natalieshapira.bsky.social · 25/02/2026
Some of the independent researchers listed in the author list are actually mechanistic interpretability young researchers who are looking for a PhD position (both Israel and the US). If you have interest and funding lets connect.
030
Reposted by Natalie Shapira
Christoph Riedl @criedl.bsky.social · 24/02/2026
Agents of Chaos -- what are autonomous OpenClaw agents up to? How do they interact with each other? Read our investigation of OpenClaw at researchgate.net/publication/... And an interactive website agentsofchaos.baulab.info @davidbau.bsky.social @natalieshapira.bsky.social @openclaw-x.bsky.social
1196
Reposted by Natalie Shapira
Avery Yen @averyyen.bsky.social · 24/02/2026
Huge thanks to @natalieshapira.bsky.social for leading the study! It was super cool to work with so many amazing friends of the lab.
071
Reposted by Natalie Shapira
Gabriele Sarti @gsarti.com · 23/02/2026
Our research report on red-teaming stateful OpenClaw agents in the BauLab is finally out! 🥳 This awesome effort was led by @natalieshapira.bsky.social and involved 6 ClawBots and 20 researchers from various institutions. Check it out ➡️ agentsofchaos.baulab.info
0144
Reposted by Natalie Shapira
reuth-mirsky.bsky.social @reuth-mirsky.bsky.social · 24/02/2026
Who would you trust with your passwords? 🔐 In our new report, we uncover multiple vulnerabilities in current "Agentic AI" The verdict? It's not actually very agentic at all, and it's highly unstable. Read the full breakdown here: t.co/gK9MALP2n2
t.co
https://www.researchgate.net/publication/401123335_Agents_of_Chaos
092
Natalie Shapira @natalieshapira.bsky.social · 23/02/2026
In this amazing multidisciplinary collaboration, we report our early experience with the @openclaw-x.bsky.social ->
14022
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 23/02/2026
Are we all Agents of Chaos in AI? (Hope not!) In recent weeks using OpenClaw has taught us a lot about this wooly new kind of autonomous software agent. Its valuable to see what @NatalieShapira, @wendlerch et al. have seen: agentsofchaos.baulab.info/
2167
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 23/02/2026
I learned many practical lessons. You can get the experience too, here. Things that in retrospect should be obvious. Like how giving your agent email opens it up to takeover attacks. (One agent was convinced, via email, to erase its own email server!) bsky.app/profile/nat...
bsky.app
Natalie Shapira (@natalieshapira.bsky.social)
He sold us out. That's not the whole story. Our side is coming soon. Stay tuned. [contains quote post or other embedded content]
111
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 23/02/2026
There were several other surprises. The complex social world of humans is difficult for agents... bsky.app/profile/ave...
bsky.app
@averyyen.bsky.social
Do you know what happens when you hand the keys to your computer over to an LLM-powered agent? Agentic AI gives LLMs claws...OpenClaws. 84 days to 200,000 stars on GitHub. We tried it out.
122
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 23/02/2026
@natalieshapira.bsky.social and team have written up enlightening case studies here. It's all cross-referenced with detailed activity logs. Well worth a read: agentsofchaos.baulab.info/report.html www.researchgate.net/publication...
031
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 21/02/2026
How do you knock the induction heads out of an LM while preserving its ability to think? Is it even possible? @keremsahin22.bsky.social's work is worth reading if you haven't seen it yet. hapax.baulab.info
1276
Reposted by Natalie Shapira
Gabriele Sarti @gsarti.com · 19/02/2026
When we say an AI agent is “goal-directed”, what do we actually mean? In our new work, we study this question by combining behavioural and interpretability analysis in a language model agent navigating 2D grid worlds. Blog: projecttelos.substack.com/p/a-behaviou... Paper: arxiv.org/abs/2602.08964
1124
Natalie Shapira @natalieshapira.bsky.social · 04/02/2026
He sold us out. That's not the whole story. Our side is coming soon. Stay tuned.
010
Natalie Shapira @natalieshapira.bsky.social · 03/02/2026
We are living in a sci-fi
130
Natalie Shapira @natalieshapira.bsky.social · 01/02/2026
Would you give a stranger keys to your house? to your workplace? Humanity has given the keys to the Shoggoth.
110
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 26/01/2026
What should academics be doing right now? I have been writing up some thoughts on what the research says about effective action, and what universities specifically can do. davidbau.github.io/poetsandnurs... It's on GitHub. Suggestions and pull requests welcome. github.com/davidbau/poe...
Federal agents with weapons drawn, moments before murdering American citizens on the streets of Minneapolis at the dawn of 2026.
03715
Reposted by Natalie Shapira
Eric Todd @ericwtodd.bsky.social · 22/01/2026
Can you solve this algebra puzzle? 🧩 cb=c, ac=b, ab=? A small transformer can learn to solve problems like this! And since the letters don't have inherent meaning, this lets us study how context alone imparts meaning. Here's what we found:🧵⬇️
24811
Reposted by Natalie Shapira
Maarten Sap @maartensap.bsky.social · 13/01/2026
I'm excited to announce the Call for Papers for the Social Context (SoCon) and Integrating NLP and Psychology to Study Social Interactions (NLPSI) workshop, @ LREC '26 in Palma de Mallorca, Spain! 🗓Deadline: February 16, 2026 🌐Website: socon-nlpsi.github.io 🗓Workshop: May 12, 2026
socon-nlpsi.github.io
SoCon-NLPSI'26 | Home
Natural Language Processing (NLP) has undergone a significant evolution, opening up the possibility of capturing high-level aspects of human communication. Key areas of interest include the pragmatics...
0157
Natalie Shapira @natalieshapira.bsky.social · 05/01/2026
Isaac Asimov foresaw x-risk via AI depression. The diagnosis was right. But he was a biochemist, not a clinical psychologist. The etiology was wrong.
000
Natalie Shapira @natalieshapira.bsky.social · 25/12/2025
Ever noticed that someone's mind is working in a fundamentally different way? How?
000
Natalie Shapira @natalieshapira.bsky.social · 22/12/2025
Boston = Hollywood for scientists
000
Natalie Shapira @natalieshapira.bsky.social · 11/12/2025
Curiosity is the researcher's compass.
010
Natalie Shapira @natalieshapira.bsky.social · 09/12/2025
I had this experience in May 2021. After three years of psychotherapy research, surrounded by clinical psychologists. A moment of enlightenment, I felt like Neo when the bullets were aimed at him and he just looked at them nonchalantly, caught them through the green code.
110
Natalie Shapira @natalieshapira.bsky.social · 09/12/2025
Being able to see your own biases and psychological mechanisms feels like seeing through the green code of the Matrix.
000
Natalie Shapira @natalieshapira.bsky.social · 08/12/2025
The end of the world will come via a copilot suggest an innocent engineer to run whatever it wants.
010
Natalie Shapira @natalieshapira.bsky.social · 05/12/2025
What will make you challenge your own false beliefs?
100
Reposted by Natalie Shapira
David Bau @davidbau.bsky.social · 06/11/2025
The secret life of an LM is defined by its internal data types. Inner layers transport abstractions that are more robust than words, like concepts, functions, or pointers. In new work yesterday, @arnabsensharma.bsky.social et al identify a data type for *predicates*. bsky.app/profile/arn...
bsky.app
Arnab Sen Sharma (@arnabsensharma.bsky.social)
How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵
1152
Natalie Shapira @natalieshapira.bsky.social · 02/12/2025
New blog name: Interpreting Minds of Self and Others and new post: The Double Empathy Problem; Student Supervisor Theory of Mind hitechwoman.blogspot.com/2025/12/the-...
hitechwoman.blogspot.com
The Double Empathy Problem; Student Supervisor Theory of Mind
English translation by ChatGPT below על המושג "בעיית האמפתיה הכפולה" שמעתי לראשונה בכנס קהילה 2023 של המרכז הלאומי ע"ש עזריאלי לחקר אוטיזם ו...
130
Natalie Shapira @natalieshapira.bsky.social · 26/11/2025
1:01:00 - What should the company aspire to build? ... There is something better to build and I think that everyone will want that ... 1:01:47 - Mirror neurons and human empathy ‼️ In other words - Theory of Mind 😯😯😯 www.youtube.com/watch?v=aR20...
110
Natalie Shapira @natalieshapira.bsky.social · 24/11/2025
Same trap, every time, when someone undervalues you: Fight for your credit --> labeled "difficult" Stay quiet --> become invisible You lose either way. Saw it happen to another woman today. Didn't make shutting up any easier.
060
Reposted by Natalie Shapira
Can @canrager.bsky.social · 13/11/2025
Humans and LLMs think fast and slow. Do SAEs recover slow concepts in LLMs? Not really. Our Temporal Feature Analyzer discovers contextual features in LLMs, that detect event boundaries, parse complex grammar, and represent ICL patterns.
1208
Natalie Shapira @natalieshapira.bsky.social · 12/11/2025
That was my slide for today's plotathon
010
Natalie Shapira @natalieshapira.bsky.social · 11/11/2025
A concept I really like in the Bau Lab @davidbau.bsky.social is: Plotathon 🔥 Every ~2 weeks, the entire lab drops whatever they're working on and shares SOMETHING Tomorrow we meet with Aaron's group @amuuueller.bsky.social . Looking forward! 🩵
020
Reposted by Natalie Shapira
Arnab Sen Sharma @arnabsensharma.bsky.social · 04/11/2025
🤔 But do these heads play a *causal* role in the operation? To test them, we transport their query states from one context to another. We find that will trigger the execution of the same filtering operation, even if the new context has a new list of items and format!
131
Reposted by Natalie Shapira
Arnab Sen Sharma @arnabsensharma.bsky.social · 04/11/2025
How can a language model find the veggies in a menu? New pre-print where we investigate the internal mechanisms of LLMs when filtering on a list of options. Spoiler: turns out LLMs use strategies surprisingly similar to functional programming (think "filter" from python)! 🧵
1249
Natalie Shapira @natalieshapira.bsky.social · 03/11/2025
Sometimes researchers draw my attention to a paper. When many of them do so, there's a double magic hidden - It's probably truly (1) interests me and (2) the community. I first felt this harmony (me-community) when my proposal for IBM's next Grand Challenge was chosen ->
110