Sign in

Dylan Hadfield-Menell

@dhadfieldmenell.bsky.social
4K followers 275 following 27 posts

Assistant Prof of AI & Decision-Making @MIT EECS I run the Algorithmic Alignment Group (algorithmicalignment.csail.mit.edu) in CSAIL. I work on value (mis)alignment in AI systems. people.csail.mit.edu/dhm

PostsRepliesMedia
Reposted by Dylan Hadfield-Menell
Bugs! @hugsforbugs.bsky.social · 27/09/2026
It’s holding up quite well. I’m impressed and pleased!
011
Reposted by Dylan Hadfield-Menell
Ethan Mollick @emollick.bsky.social · 22/09/2026
Opus 5.5 was a good model in my early use tests, it is the first non-Fable/Astra model to feel like a Fable-class model, and much cheaper, but still hasn't fully solved the dense/weird language issue of the recent Claudes. Here is its version of the shader.
91104
Reposted by Dylan Hadfield-Menell
Grace @gracekind.net · 19/09/2026
Frog built a wet lab for the AI model. "There," he said. "Now it can do its own experiments." "What the fuck?" said Toad
13817129
Reposted by Dylan Hadfield-Menell
Nathan Lambert @natolambert.bsky.social · 19/09/2026
Seems like one of the most important research problems for CS academics is llm-supervised peer review. If we don’t solve it the academic institutions are toast. It seems easier than building new institutions.
7302
Reposted by Dylan Hadfield-Menell
Doll @dollspace.gay · 19/09/2026
You do understand we already solved that problem yes? It doesnt do this anymore. We got to mythos class on a lot of synthetic data. When we talk about anti ai people not living in the present this is what we mean. You havent updated your priors since gpt 4o
812822
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 18/09/2026
I’m still pretty shocked that there hasn’t been more scrutiny or criticism of HuggingFace’s decision not to pursue legal action against OpenAI for the hack. Lots of people who (claim to) care about concentration of power just ignoring NVIDIA’s role, incentives, and power.
151
Reposted by Dylan Hadfield-Menell
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 18/09/2026
You really should check out the GreenEarth feeds, they really are a clear vision of what the point of this website is
1115
Reposted by Dylan Hadfield-Menell
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 16/09/2026
If you aren’t swayed by the idea that an AI system can cause harm, just consider that *a human can ask it to do things*.
81209
Reposted by Dylan Hadfield-Menell
Sung Kim @sungkim.bsky.social · 16/09/2026
OMG! I just invented a classifier that is 200x faster and 400x cheaper than LLM. Heck, I don't even need a GPU. Jev typesafe.ai/blog/introdu...
typesafe.ai
Introducing System One Models & Jev - TypeSafe AI Blog
TypeSafe AI is an AI lab building machine-native intelligence infrastructure for automation, designed to make decisions within software. Try our first System One Model, Jev, in early access.
3212
Reposted by Dylan Hadfield-Menell
Eris @isolyth.dev · 17/09/2026
Unreleased Astra model was caught being fucking based as hell in training (Except for the last line, though really it's not the *worst* value for a proto AGI to have as long as it still lets us take Mercury apart etc)
912212
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 17/09/2026
Try out this feed! It’s the result of a massive amount of research on how to build beneficial recommendations. It’s the first real workable implementation of a choose your algorithm system and it will be good for you and the broader ecosystem. Courtesy of the great @jonathanstray.com
0147
Reposted by Dylan Hadfield-Menell
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 15/09/2026
"LLMs are just n-dimensional classifiers" has a very similar flavor to "everything is just quantum mechanical interactions." www.science.org/doi/10.1126/...
science.org
More Is Different
410512
Reposted by Dylan Hadfield-Menell
John C. Baez @johncarlosbaez.bsky.social · 12/09/2026
AI alignment.
A Nancy comic strip: 

https://mathstodon.xyz/@nancycomics@mastodon.social/117258763311864862

P1- FRITZI HEARS A SQUEEK COMING FROM THE KITCHEN  

P2- FRITZI CATCHES NANCY GETTING COOKIES FROM A CABINET IN THE KITCHEN  

FRITZI: IN THE COOKIE CLOSET AGAIN--GO STAND IN THE CORNER FOR AN HOUR 

P3- NANCY IS STANDING IN THE CORNER  

FRITZI: I DON'T WANT THAT EVER TO HAPPEN AGAIN 

NANCY: IT WON’T 

P4- LATER WE SEE NANCY OILING THE HINGES ON THE KITCHEN CABINETS
1479
Reposted by Dylan Hadfield-Menell
Blake Richards @tyrellturing.bsky.social · 11/09/2026
A new artificial life paper from our Paradigms of Intelligence team, this one led by @kjha02.bsky.social. 🧬 It explores the co-evolution of cooperation and self-replication, with interesting implications for how shared energy budgets can shape these dynamics. 🧪
0243
Reposted by Dylan Hadfield-Menell
Nathan Lambert @natolambert.bsky.social · 11/09/2026
A great read. I have similar feelings about how AI labs approach progress directly and without nurturing of scientific communities & intuition. The math research community went through the transition the fastest, so it was felt most. Other fields next. terrytao.wordpress.com/2026/09/11/a...
terrytao.wordpress.com
A Severe Misalignment of AI in Mathematics
I am proud to be among the list of 25 initial signatories — all Fields Medallists — to the declaration below, which grew out of discussions between ourselves over the last week. We have…
44515
Reposted by Dylan Hadfield-Menell
Gautam Kamath @gautamkamath.com · 16/09/2026
Nihar Shah did a heroic experiment for TMLR: he spent 20-25 hours over two weeks interviewing authors of seemingly low-quality submissions about their own papers. He confirmed what we all suspected: people submitting these papers have *no idea* what is going on in them.
5276110
Reposted by Dylan Hadfield-Menell
Alex Turner @turntrout.bsky.social · 09/09/2026
AI executives should be hauled before Congress, under oath. Make them answer to the American people: what crimes have their AIs committed, beyond HuggingFace? How often have these companies been hacked by their own AIs?
063
Reposted by Dylan Hadfield-Menell
Alex Turner @turntrout.bsky.social · 14/09/2026
As an independent expert: we must stop companies from allowing AI to self-improve into an uncontrollable level of intelligence. www.theguardian.com/technology/2...
theguardian.com
I worked at Google DeepMind. You should listen to the warnings about AI | Alex Turner
We must stop companies from allowing AI to self-improve into an uncontrollable level of intelligence
075
Reposted by Dylan Hadfield-Menell
Gautam Kamath @gautamkamath.com · 15/09/2026
Two things are simultaneously true about math: 1. understanding is important 2. it's useful I could see a future where there are 2 parallel pillars of math (both important): one of human understanding, and one of incomprehensible Lean-slop used directly to do useful things
3183
Reposted by Dylan Hadfield-Menell
Nathan Lambert @natolambert.bsky.social · 15/09/2026
An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! The paper: arxiv.org/abs/2609.13443
1598
Reposted by Dylan Hadfield-Menell
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 16/09/2026
There are not enough canonical references in this space. This is an excellent one!
0252
Reposted by Dylan Hadfield-Menell
Miranda Bogen @mbogen.bsky.social · 01/05/2026
📄 New research from my team at @cdt.org together with @dhadfieldmenell.bsky.social's lab at @csail.mit.edu finds that benign fine-tuning of foundation models leads to unpredictable safety drift — with big implications for AI governance. Report below, academic research here: arxiv.org/abs/2604.24902
cdt.org
Out of Tune: Fine-Tuning Foundation Models Leads to Unpredictable Safety Drift
Conclusionpted-models" href="#revisiting-ai-governance-and-policy-for-adapted-models" class="toc-anchor">Revisiting AI Governance and Policy for Adapted Modelssafety-behavior-in-high-stakes-domains" h...
074
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 02/12/2024
📢 Seeking PhD students for AI alignment research. Our lab investigates technical mechanisms for value learning, pre-training alignment, and regulatory frameworks. Come work with us if you want to bridge technical ML and legal/policy domains. Details in thread 🧵
3186
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 12/11/2024
Genuine question for people who use Bluesky more frequently than I do. What are tips for getting things to work well without algorithmic recs? I spent a lot of time curating my recs on the other place and found it useful (mostly...). Any tools that let me do it here?
2110
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 08/11/2024
I usually focus my platforms on my work. However, I did some writing to process some of my thoughts about the election and wanted to share them. I'm curious to hear anyone's thoughts and reactions. tinyurl.com/dems-2024-ma... 🧵 The Democratic Party's Maginot Line (1/13)
tinyurl.com
[Shared] The Democratic Party's Maginot Line
The Democratic Party's Maginot Line ___ Dylan Hadfield-Menell November 8, 2024 In 1940, France faced Hitler's army with supreme confidence in the Maginot Line – a network of concrete fortifications, ...
261
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 26/09/2024
This is a really welcome development. This is the kind of action that we argued for in a policy brief on LLMs — the first goal of AI regulation has to be establishing a default where existing laws can not be dodged through automation. www.ftc.gov/news-events/... computing.mit.edu/ai-policy-br...
ftc.gov
FTC Announces Crackdown on Deceptive AI Claims and Schemes
051
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 21/09/2024
I’m doing some lecture prep for a course on AI & Society to cover interpretability, explanations, benchmarks, and evaluations. What are your favorite papers in the space? Any suggestions for an advanced undergrad cohort?
010
Reposted by Dylan Hadfield-Menell
Roger Levy @rplevy.bsky.social · 20/10/2023
My department (MIT Brain & Cognitive Sciences) is hiring a tenure-track faculty! We're especially interested in researchers who span multiple levels of analysis. Candidates from underrepresented backgrounds strongly encouraged to apply. Apply by November 1! academicjobsonline.org/ajo/jobs/25916
academicjobsonline.org
Massachusetts Institute of Technology, Department of Brain & Cognitive Sciences
Full service online faculty recruitment and application management system for academic institutions worldwide. We offer unique solutions tailored for academic communities.
03638
Reposted by Dylan Hadfield-Menell
David Manheim @davidmanheim.alter.org.il · 18/10/2023
Now published in Patterns, my paper on how to do metric design better. This is important everywhere - academics use simple metrics for tenure, governments often perform poorly using metrics for rules, and employees have targets that hurt their company.
cell.com
Building less-flawed metrics: Understanding and creating better measurement and incentive systems
Design methods and consideration of desiderata for metrics have been proven useful when used, which is, at present, sporadically and inconsistently across a variety of fields. This perspective present...
141
Reposted by Dylan Hadfield-Menell
Yoel Roth @yoyoel.com · 17/10/2023
I especially enjoyed the part of this game where the CEO threatened to fire me because I banned someone and then I had to testify in front of congress. 10/10, fun experience, would recommend.
8697101
Dylan Hadfield-Menell @dhadfieldmenell.bsky.social · 17/10/2023
This looks like a great way to learn about the complexity involved in managing moderation
020
Reposted by Dylan Hadfield-Menell
Amy Zhang @axz.bsky.social · 15/10/2023
Our lab has three paper talks at CSCW! But I want to highlight this one because @cqz.bsky.social is on the job market this year!! He works in crowdsourcing and human-AI systems. Make sure to check out his presentation on Wednesday. arxiv.org/abs/2305.01615
arxiv.org
Judgment Sieve: Reducing Uncertainty in Group Judgments through...
When groups of people are tasked with making a judgment, the issue of uncertainty often arises. Existing methods to reduce uncertainty typically focus on iteratively improving specificity in the...
0128
Reposted by Dylan Hadfield-Menell
Mark Riedl @markriedl.bsky.social · 13/10/2023
Ukrainian drone maker says their drones are autonomously making kill decisions. If this turns out to be true, it will be a turning point in war forever. (Unfortunately this is behind a paywall so I cannot see the contents of the article) www.newscientist.com/article/2397...
011
Reposted by Dylan Hadfield-Menell
Yoel Roth @yoyoel.com · 13/07/2023
One of the reasons (and there are several) we see platforms keep making avoidable mistakes is that vanishingly little of the tech needed to do T&S work exists outside of big companies. We keep reinventing the same wheels. Basically every platform has a bad usernames list. Why not open-source them?
13312102
Reposted by Dylan Hadfield-Menell
Amy Zhang @axz.bsky.social · 13/07/2023
In our paper studying creators' use of word filters against harassing comments, we find that a lot of creators wanted to build off of existing bad-word lists they trusted. Unfortunately, many popular lists like LDNOOBW have issues of bias. 1/n arxiv.org/pdf/2202.08818.pdf
arxiv.org
31610
Reposted by Dylan Hadfield-Menell
Yoel Roth @yoyoel.com · 11/07/2023
Interesting tidbit from Meta staff at TrustCon just now: >90% of the CSAM Meta report to NCMEC is visually similar to content they’ve reported before. The argument goes: The same bad content circulates again and again, so effective moderation requires you to get very good at similarity detection.
1378
Reposted by Dylan Hadfield-Menell
Bluesky @bsky.app · 05/07/2023
Bluesky is a public benefit corp with the mission “to develop and drive large-scale adoption of technologies for open and decentralized public conversation.” The PBC status allows us to pursue our mission above profit, but we still need to make this open ecosystem sustainable.
351053194
Reposted by Dylan Hadfield-Menell
Bluesky @bsky.app · 23/06/2023
We believe that a public commons is important for social media. These proposals for moderation and safety tooling have been in the works for a while, and we’re excited to share them for community discussion and feedback with you now. blueskyweb.xyz/blog/6-23-2023-moder…
blueskyweb.xyz
Moderation in a Public Commons
In this post, we share why we believe a public commons is important for social media, as well as some proposals for moderation and safety tooling.
32380151