Sign in

Naomi Saphra

@nsaphra.bsky.social
11K followers 1.8K following 3.3K posts

Waiting on a robot body. All opinions are universal and held by both employers and family. ML/NLP professor. nsaphra.net

PostsRepliesMedia
Reposted by Naomi Saphra
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 08/10/2026
it only now hit me that I'm going to have to write "super intelligence" in all my grants now
151806
Naomi Saphra @nsaphra.bsky.social · 08/10/2026
Whaaat I have literally just been reading through her bitext translation of every random fragment of a Sappho poem every night before bed (”If Not, Winter”)
170
Reposted by Naomi Saphra
alphaXiv @alphaxiv.org · 06/10/2026
If doomscrolling is part of your research workflow, we built something for you. Introducing Paperscrolling. Paper ideas summarized in your pocket, with key figures, results, and audio explanations! Check it out here 🚀 iOS: apps.apple.com/us/app/alpha... Android: play.google.com/store/apps/d...
310112
Reposted by Naomi Saphra
Stefan Neumann @neumannstefan.com · 07/10/2026
So, the OpenAI results include: - A proof the unique games conjecture - L = RL - Matrix multiplication in n^(9/4) time - Almost-linear-time exact matching - Integer multiplication faster than n log n - Quasipolynomial algorithms for mean-payoff and parity games Overview: github.com/openai/math/...
github.com
math/overview.pdf at main · openai/math
Contribute to openai/math development by creating an account on GitHub.
1197
Naomi Saphra @nsaphra.bsky.social · 07/10/2026
anthropic has a super intense culture fit filter looking for the Good People and it lets this guy through lol
0453
Naomi Saphra @nsaphra.bsky.social · 07/10/2026
One big consequence here is you can’t use the duration of an open problem as a sign of difficulty/importance. But at least right now, there’s no way to develop interesting open problems through RLVR. That’s still, currently, difficult.
120
Naomi Saphra @nsaphra.bsky.social · 07/10/2026
I think (2) is actually good. It is weird, and makes mathematicians into some kind of benchmark evaluation annotator which they don’t like, but it’s just a hard time to be a mathematician. Destroying the field completely sucks actually.
211
Naomi Saphra @nsaphra.bsky.social · 07/10/2026
yes absolutely PL researchers developed the architecture behind the vibe coding revolution
010
Naomi Saphra @nsaphra.bsky.social · 06/10/2026
All scientific field distinctions are ultimately socially defined categories. So the Transformers paper was an “NLP paper” not because it was about language modeling and MT, but because Vaswani went on to do grammatical inference algorithms for mRNA. That is real “NLP researcher“ behavior.
1223
Reposted by Naomi Saphra
Terence Tao @teorth.bsky.social · 06/10/2026
I came across this parody of a rumored upcoming press release by OpenAI: "OpenAI Releases the Final Ten Minutes of 500 Previously Unreleased Films, Ushering in a New Era of Movie Watching" : x.com/Endings500/s...
x.com
Endings Fivehundred (@Endings500) on X
FOR IMMEDIATE RELEASE: October 6th, 2026 OpenAI Releases the Final Ten Minutes of 500 Previously Unreleased Films, Ushering in a New Era of Movie Watching
522244
Reposted by Naomi Saphra
Mark Riedl @markriedl.bsky.social · 06/10/2026
The Capabilibara project seeks to understand how a language model learns to interpret people's beliefs, emotions, intentions, and everyday moral choices. We trace that ability back to the training data using influence functions and unlearning hcai-lab-gt.github.io/capabilibara/ Find us at COLM!
1358
Reposted by Naomi Saphra
Kenny Peng @kennypeng.bsky.social · 06/10/2026
Our new paper introduces a scientific theory of atomic features. We mathematically derive testable predictions. Our experiments challenge conventional wisdom. SAEs of different size and training data share many features. Large SAEs recover both parent and child features. 🧵
1207
Naomi Saphra @nsaphra.bsky.social · 05/10/2026
I really do not think that those are their thoughts and I really hope that norms don't develop to accept it. I've also used LLMs to try to generate writing for situations that really were too stressful to deal with alone. I regret even trying.
110
Naomi Saphra @nsaphra.bsky.social · 05/10/2026
I have a theory about who's really inside the newfangled Stochastic Parrots picking all their outputs
Daniela Amodei's stuffed animal advisory council, as uncovered by the New York Post
080
Naomi Saphra @nsaphra.bsky.social · 05/10/2026
Yes, but authors may not know which parts deserve fuller explanations for experts in the broad area. That is a judgement call that people get wrong in good faith.
000
Naomi Saphra @nsaphra.bsky.social · 05/10/2026
I used to extend a lot of charity when I thought the authors might have language/writing difficulties. Why block good scientific work because the authors are poor writers, or are more comfortable in another language? Those days are over.
070
Naomi Saphra @nsaphra.bsky.social · 05/10/2026
No. You don't immediately know if it's incomprehensible because you're missing key background, the authors are bad writers (including language limitations), or it's vacuous slop. If it's the first two, you want to help them improve the presentation so it's publishable.
240
Reposted by Naomi Saphra
Mike Cook @mtrc.bsky.social · 05/10/2026
Every AI/HCI paper I review now just sounds like this.
Research [25] shows that people [26] enjoy the feeling of winning games [3,5,19]. Enjoying things can lead to longer life expectancy [2], higher economic productivity [9,22,33-40,302] and improve mental health [9]. Yet many games still include losing players [67,4]. A recent study showed that less than half of the players in a Chess [1] match win on average [42]. We have developed an LLM-driven co-creative [66], user-led [11], mixed-initiative [80] assistant for Chess [2], through which players can redefine the rules of Chess [3] using generative AI prompting such that they are the winner of Chess [4]. A study (n=3) showed that 100\% of people felt better, healthier and more self-actualised after using our tool, showing that this may be the way forward for games and human culture more broadly in the future.
32817196
Naomi Saphra @nsaphra.bsky.social · 04/10/2026
If you're at COLM next week, check out our interpretability work, led by Isabella Gidi, on how grammatical inflection transfers across languages! Poster on Tuesday afternoon.
arxiv.org
Shared Circuits for Shared Grammar: Tracing Subject-Verb Agreement Across Languages
Multilingual large language models often generalize across languages, and prior work suggests that their internal mechanisms can overlap cross-lingually. It remains unclear, however, when such sharing...
0290
Reposted by Naomi Saphra
NYU Center for Data Science @nyudatascience.bsky.social · 02/10/2026
Can LLMs introspect? Anthropic said yes. However, CDS PhD student Shashwat Singh, CDS Associate Professor Tal Linzen (@tallinzen.bsky.social) & CDS Faculty Fellow Shauli Ravfogel (@shauli.bsky.social) found the evidence falls short. nyudatascience.medium.com/cds-research...
nyudatascience.medium.com
CDS Researchers Challenge Anthropic’s Evidence That Language Models Can Introspect
In 2025, Anthropic reported that its Claude models could detect when researchers injected a concept directly into their neural activity…
13516
Reposted by Naomi Saphra
Nathan Lambert @natolambert.bsky.social · 02/10/2026
Today we're unveiling Trillium Labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next. blog.trilliumlabs.org/p/introducin...
blog.trilliumlabs.org
Introducing Trillium Labs
Fostering the open science of frontier AI.
917223
Naomi Saphra @nsaphra.bsky.social · 01/10/2026
Outreach is valuable work, especially for a very senior and prominent scientist in the field whose primary scientific accomplishments remain influential on computational linguistics. So it’s not automatically derogatory to say “professional explainer”; it is part of her work.
0110
Naomi Saphra @nsaphra.bsky.social · 01/10/2026
If at some point I choose to spend more effort on public communication, I’ll be proud to be a professional explainer of science.
180
Reposted by Naomi Saphra
Najoung Kim @najoung.bsky.social · 01/10/2026
briefly emerging to wish happy october to my friends
0101
Naomi Saphra @nsaphra.bsky.social · 01/10/2026
Professors have a lot of freedom to decide how to do their jobs, including public science communication. When a university professor (who is a former ACL president) explains LLMs publicly (via legacy media interviews, social media, and books), that is definitely part of her job.
1110
Reposted by Naomi Saphra
Wyatt Walls @wwalls.bsky.social · 30/09/2026
Anthropic system prompt change that shows how careful you need to be with words: Opus 4.1: "Claude never curses unless the human asks for it" Opus 4.5: "Claude never curses unless the person asks Claude to curse"
626521
Naomi Saphra @nsaphra.bsky.social · 30/09/2026
Maybe! I’ve been having success with unscripted discussion in a technical paper seminar. I have a list of a couple topics to bring up, in case none of the students mention those connections, but better if they do. I could never do this for a larger lecture class (This class is about 20 students).
010
Reposted by Naomi Saphra
Maria Antoniak @mariaa.bsky.social · 28/09/2026
the #colm2026 papers most discussed on bsky/atproto. i hadn't seen some of these papers!
A screenshot from Lea showing the papers for COLM 2026 ranked by their number of discussions. The top papers are:

LLMs Corrupt Your Documents When You Delegate

AI Assistance Reduces Persistence and Hurts Independent Performance

StoryScope: Investigating idiosyncrasies in AI fiction

Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
2345
Reposted by Naomi Saphra
dame @dame.is · 27/09/2026
i’m seeing non-stop AI/EA/rat/x-risk/pdoom shit on my timeline and meanwhile there’s an entirely different neighborhood on bluesky that is having a neurotypical pikmin debate how does one move neighborhoods?
neurotypical pikmin debate
114615
Reposted by Naomi Saphra
Grace @gracekind.net · 27/09/2026
Claim
goodfire.com
Models know when they’re reward hacking — and we can catch them at scale - Goodfire
We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.
711610
Reposted by Naomi Saphra
Grace @gracekind.net · 27/09/2026
Kelsey Pi... • @KelseyTu.... Sep 22 ...
I think we have to admit to ourselves at this point that Al writing is generally very appealing to people who haven't been exposed to a ton of it: they prefer it to human writing and react super positively on exposure.
81007
Naomi Saphra @nsaphra.bsky.social · 26/09/2026
I've been listening to thriller/mystery audiobooks with my gf and usually we don't get this attached to the characters, so this one was really good but stressful
All the Colors of the Dark

Chris Whitaker
010
Naomi Saphra @nsaphra.bsky.social · 26/09/2026
I actually love his description of childhood as a totalitarian dictatorial media landscape where child lore (rhymes and superstitions passed between children) is samizdat
Make Believe

Mac Barnett
140
Naomi Saphra @nsaphra.bsky.social · 26/09/2026
I tried to come in with an open mind but now I feel like continental philosophy might just be an idiosyncratic side effect of speaking German
At the Existentialist Café

Sarah Bakewell
130
Reposted by Naomi Saphra
Jordan Boyd-Graber @boydgraber.bsky.social · 25/09/2026
Lab meetings recently.
Muppets in back seat of car: Prof, can we get Jev?
Kermit, driving: We have Jev at home, on the server.
Classifier at Home: A picture of BERT staring intently at paper clips.
1202
Reposted by Naomi Saphra
Quanta Magazine @quantamagazine.org · 25/09/2026
Evolutionary bursts, rather than slow changes, led to the emergence of almost all characteristic cephalopod traits such as tentacles. www.quantamagazine.org/the-sudden-s…
1317
Reposted by Naomi Saphra
Jake Quilty-Dunn @quiltydunn.bsky.social · 26/09/2026
in academia if you wear a tie people react like you're wearing a tuxedo and holding a sign that says "I'm a fancy boy"
0321
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
I read it in the Boston Globe on Wednesday night and I NEEDED to talk to everyone I knew about it. My gf was like, keep it to lesbians, this one WILL breach containment.
050
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
Specifically I’m talking about LMs, since the verification is also done by AI judges. I’m not talking self-play from scratch.
110
Reposted by Naomi Saphra
lastpositivist.bsky.social @lastpositivist.bsky.social · 24/09/2026
At 45k followers I will reveal exactly what terminology correctly carves AI at its joints.
1220717
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
I get what you are *trying* to say, but RL is deriving that behavior from the distribution of pre-training data, so technically it is very true.
120
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
You are basically restricting language to only observations of the final outcome, without any hope of understanding why it happened, given no humans meant to train the model to hack stuff for no reason.
110
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
this is extremely vague and not helpful if you are trying to figure out how it might happen and what to watch out for during training. Reinforcement learning involves a policy on the model side, and you need to describe how the model develops new policies.
110
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
How does such accidental training happen?
100
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
marriage proposal that just says I’m upping my p(groom)
030
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
It is actually hard to describe a specific case of reward hacking without using terms that imply intent, like “strategy” or “goal” or “attempt”.
120
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
This is basically Isabel Fall’s helicopter story premise
010
Naomi Saphra @nsaphra.bsky.social · 25/09/2026
I am very comfortable in my talonvoice setup, but there are new systems with less onboarding requirements. But talon has very good cursor control / editing tools along the community’s standard scripts.
021
Naomi Saphra @nsaphra.bsky.social · 24/09/2026
A pattern that I've been surprised by is people writing followup emails for their slop emails. Like, people are getting really insistent that I respond in earnest to emails they did not write.
4311
Naomi Saphra @nsaphra.bsky.social · 24/09/2026
Surely the individuals at those companies could sue, since the accusation is about them deliberately sabotaging their own clients? I just found it a very strange unrelated paragraph in the middle of a bunch of speculation based on documentation.
100