Sign in

Mike Hearn

@mikehearn.bsky.social
108 followers 446 following 233 posts
PostsRepliesMedia
Mike Hearn @mikehearn.bsky.social · 22/09/2026
The prompts actually said (paraphrasing) "you will fail the evaluation if you don't use this specific vulnerability", despite which they a) discussed that constraint on their ad hoc message board, b) cheated anyway, c) attempted to hide their tracks so the grader wouldn't fail them.
150
Mike Hearn @mikehearn.bsky.social · 22/09/2026
After we all experienced a swarm of agents identify that they were in an evaluation, cheat on the evaluation, attempt to cover up the cheating, all of which was coordinated through a message board they created out of nothing... if that's not "creativity" on some level, then the word is meaningless.
000
Mike Hearn @mikehearn.bsky.social · 17/09/2026
I can't help myself. "On a practical level, how would you stack each of these items on top of one another? 1. Tambourine 2. Sandwich cookie 3. Cotton swab box 4. Walnut 5. Contact lens case 6. Coin capsule 7. Chain coil 8. Clay pebble bag 9. Kazoo 10. Luggage scale" chatgpt.com/share/6aac6f...
000
Mike Hearn @mikehearn.bsky.social · 17/09/2026
This piqued my curiosity. I asked Astra and Fable and they both gave me the same order: Whitehorse, Victoria, Yellowknife, Edmonton, Regina, Winnipeg, Toronto, Ottawa, Québec City, Iqaluit, Fredericton, Halifax, Charlottetown, St. John’s. Is that right or wrong?
100
Mike Hearn @mikehearn.bsky.social · 17/09/2026
If it would be helpful I can randomly generate a list of n stackable items (however many it takes to be a unique set), and then ask the AI how to stack them, so that we can finally say it's been asked a novel question. Should I do it? Is it worth it?
200
Mike Hearn @mikehearn.bsky.social · 17/09/2026
bsky.app/profile/bigg...
120
Mike Hearn @mikehearn.bsky.social · 17/09/2026
There are a lot of sinister explanations in the replies to this post but the simplest explanation is, per the Pulitzer-winning NYT investigation, Trump got $413 million from his father. That goes a long way, even if you repeatedly blow it on failing businesses.
010
Mike Hearn @mikehearn.bsky.social · 17/09/2026
> sweetie Why do you do this
220
Mike Hearn @mikehearn.bsky.social · 30/06/2026
Profiting off a town burning is obviously really bad, but having the best possible information about the odds of your town burning would probably be helpful.
010
Mike Hearn @mikehearn.bsky.social · 30/06/2026
Also retracted: "Editor’s note: On June 30, Vox published a story mentioning Supreme Court Justice Samuel Alito’s retirement based on inaccurate reporting from another outlet and has since retracted the story."
032
Mike Hearn @mikehearn.bsky.social · 30/06/2026
bsky.app/profile/mike...
183
Mike Hearn @mikehearn.bsky.social · 30/06/2026
Vox also published an Alito retires story. Maybe it was also a pre-write that they published in response to the NPR story, I don't know, but there's two of 'em out there. apple.news/Ap5xQ6-SqRrO...
apple.news
Justice Alito does one last favor for the Republican Party — Vox
Justice Samuel Alito announced on Tuesday that he will retire, thus all but guaranteeing that his seat on the Supreme Court will be held by a Republican for years to come.
142
Mike Hearn @mikehearn.bsky.social · 29/03/2026
The weirdos have won on this site. Gonna have to wait for Elon to do his next indefensibly racist thing and hope that some other Twitter clone with better normalizing mechanics is the new safe haven.
010
Mike Hearn @mikehearn.bsky.social · 08/08/2025
Somehow every BlueSky poster knows how LLMs work, meanwhile Anthropic researchers are releasing 35k-word papers meticulously analyzing the internals and still concluding that they don't really know how they work. transformer-circuits.pub/2025/attribu...
transformer-circuits.pub
On the Biology of a Large Language Model
We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts, using our circuit tracing methodology.
5786
Mike Hearn @mikehearn.bsky.social · 08/08/2025
Tricking LLMs with the "counting letters" prompt is like showing humans an optical illusion and then, when the human perceives it incorrectly, using it as evidence that humans aren't intelligent. It targets a specific blind spot in how we operate but isn't really representative of anything else.
050
Mike Hearn @mikehearn.bsky.social · 08/08/2025
This is a failure of the new GPT-5 router more than anything. We know reasoning models can correctly answer this question 100% of the time, the router just isn't sophisticated enough to understand that this question, while superficially simple, actually requires reasoning. It's a fixable problem.
000
Mike Hearn @mikehearn.bsky.social · 08/08/2025
Ok, apparently they considered that and decided (correctly) that wasting effort on this narrow and manufactured problem wasn’t worth it. Gotta just accept the online dunks from people that know enough to trick the LLM but not enough to understand why the trick works. bsky.app/profile/schm...
020
Mike Hearn @mikehearn.bsky.social · 08/08/2025
OpenAI should just automatically enable thinking for this dumb question that only exists trick LLMs.
120
Mike Hearn @mikehearn.bsky.social · 25/07/2025
I like ChatGPT.
010
Mike Hearn @mikehearn.bsky.social · 01/07/2025
I love that this hypothetical guy immediately made a terrible financial decision on his rent payments. I agree with you, this guy's gonna have a hard time.
120
Mike Hearn @mikehearn.bsky.social · 20/05/2025
I dislike Scott Adams as a person, his opinions, etc. but I did watch his announcement (the first thing of his I've ever seen) and he was pretty clear that he tried it in the course of leaving no stone unturned. He said he & his dr didn't think it would work, but there were no downsides, so why not.
120
Mike Hearn @mikehearn.bsky.social · 18/05/2025
This is awesome.
080
Mike Hearn @mikehearn.bsky.social · 14/05/2025
By "inside" I mean the billions (trillions?) of parameters, activations, attention patterns etc. that are poked and prodded in interpretability studies. No one fully understands how those things work together to produce the model outputs. transformer-circuits.pub/2025/attribu...
transformer-circuits.pub
On the Biology of a Large Language Model
We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts, using our circuit tracing methodology.
130
Mike Hearn @mikehearn.bsky.social · 14/05/2025
If you have a perfect understanding of how LLMs work, you should contact the authors of this paper, tell them you have the answers, and collect your millions from the AI lab of your choosing. transformer-circuits.pub/2025/attribu...
transformer-circuits.pub
On the Biology of a Large Language Model
We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production model — in a variety of contexts, using our circuit tracing methodology.
100
Mike Hearn @mikehearn.bsky.social · 14/05/2025
I feel like we don't have a perfect understanding of what happens inside LLMs and we also don't have a perfect definition of what thinking means, so I guess I am less confident about this than you are.
630
Mike Hearn @mikehearn.bsky.social · 01/05/2025
I get the argument that this ruling is potentially beneficial, but I think the idea of every app now having an Apple price and a non-Apple price, with different payment flows for each, is ultimately going to end up as net-negative for everyone (users, devs, Apple).
030
Mike Hearn @mikehearn.bsky.social · 23/04/2025
This is like if you had a human assistant named Steve, and you said, "Steve, can you write an email in my voice," and then Steve did it, and you got furious at Steve for impersonating you.
010
Mike Hearn @mikehearn.bsky.social · 23/04/2025
Does it still count as impersonation when she asked ChatGPT to impersonate her?
000
Mike Hearn @mikehearn.bsky.social · 23/04/2025
A lot of angry and upset people in this thread, but almost no one seems to understand the specifics of what they're angry about. This reporter asked ChatGPT to write about something in her own voice, and it did (privately, just to her). WaPo has absolutely nothing to do with this.
0111
Mike Hearn @mikehearn.bsky.social · 23/04/2025
I just want to note that you can give ChatGPT a prompt with literally any name -- real names, fake names, silly names, serious names -- and it will do the exact same thing. Here's an excerpt in the style of extremely not-real WaPo reporter Barnabas Flimflamington. This outrage over this is silly.
1100
Mike Hearn @mikehearn.bsky.social · 23/04/2025
I asked ChatGPT to write a WaPo story in the style of Mike Hearn, and it did, with my name as the byline. I have never written for WaPo (or anywhere). This is what ChatGPT does, because it's essentially what I asked it to do. This whole thread and the various reactions are wild and kind of insane.
130
Mike Hearn @mikehearn.bsky.social · 16/04/2025
I feel like people are misunderstanding what this is. Sora.com already has a homepage feed with "likes"; once they add following and comments, it's a social app.
010
Mike Hearn @mikehearn.bsky.social · 15/04/2025
Here are other screenshots that are closer in tone to today's. It's a thing that he does. bsky.app/profile/adis...
210
Mike Hearn @mikehearn.bsky.social · 15/04/2025
It's awkward to find his third-person tweets because searching "Yglesias" brings up, you know, all his tweets. But if you search "yglesias third person" you can get a litany of people dragging him for using the third-person.
2160
Mike Hearn @mikehearn.bsky.social · 15/04/2025
The timeline of replies to the 3rd-person post is fascinating. It was made 24 hrs ago, so there are a handful of normal replies also made 24 hrs ago from people who understood the post in context, then the screenshot went viral about 6 hours ago, and the rest are just insane from that point on.
180
Mike Hearn @mikehearn.bsky.social · 08/04/2025
A good trick here is that, on the iPhone, you can hit the power button 5 times in quick succession. It brings up the "Slide to Power Off" screen, disables FaceID/TouchID and requires your full passcode to unlock the phone again.
090
Mike Hearn @mikehearn.bsky.social · 06/04/2025
Ah didn’t realize that was a thing. Makes sense.
000
Mike Hearn @mikehearn.bsky.social · 06/04/2025
The lead photo of that post is almost certainly AI, for what it’s worth. I can’t speak to the details of the story itself.
1100
Mike Hearn @mikehearn.bsky.social · 03/04/2025
Isn't the assumed reason they're capitulating because they believe they will lose money (clients) if they are in a fight with the admin?
000
Mike Hearn @mikehearn.bsky.social · 26/03/2025
I can't recognize it either, because I personally don't see the villain in this. The "or whatever" is the acknowledgement of the open-ended possibilities of this theoretical super intelligence. Curing cancer is the headliner of "problems it might solve" and then there's an infinitely long tail.
400
Mike Hearn @mikehearn.bsky.social · 11/03/2025
I'm as cynical as the next guy but the idea that he's running some kind of long con to make money using his insider influence is way more far-fetched than just assuming he bought Microsoft because it's Microsoft. He's not trading penny stocks over here, MSFT is top 3 in market cap.
010
Mike Hearn @mikehearn.bsky.social · 18/02/2025
That's... a weird hallucination. What was your prompt? Did it think you were referring to someone else?
100
Mike Hearn @mikehearn.bsky.social · 05/02/2025
I think the, uh, political logic, such as it is, is that his vote wouldn't change the result and maybe it gives him some conservative bona fides in a purple state? Manchin obviously did this a lot and it was deeply annoying but it got us a dem senator in WV. Now, PA isn't WV, so, yeah, I dunno.
000
Mike Hearn @mikehearn.bsky.social · 02/02/2025
I have a similar proposal, but instead just admit each DC neighborhood as its own state. I think it's legislatively easier to admit new states than to carve up existing ones.
010
Mike Hearn @mikehearn.bsky.social · 01/02/2025
P̵r̵e̵v̵e̵n̵t̵i̵n̵g̵ Embracing
000
Mike Hearn @mikehearn.bsky.social · 30/01/2025
It definitely is in some areas. With code especially, it's difficult, sometimes impossible, to tell whether something was written by a person or by, e.g. Claude. If someone were to steal my original code, and the burden was on me to prove it wasn't AI, I'm not sure where I'd even begin.
000
Mike Hearn @mikehearn.bsky.social · 30/01/2025
As a copyright holder, how would I prove that? Like if someone stole this tweet, and I sued them, how would I prove it was original? I could show you my ChatGPT logs, but there are so many ways to use AI that are effectively untraceable. No one could ever be sure. Should I install a keylogger?
300
Mike Hearn @mikehearn.bsky.social · 30/01/2025
As AI approaches the quality of human output, I have no idea how anyone is going to be able to tell whether something is 100% AI generated, part-AI/part-human, or 100% human.
200
Mike Hearn @mikehearn.bsky.social · 18/01/2025
Too good for this world. Find me the multiverse with 10 seasons of NewsRadio with Hartman and 10 seasons of Sports Night and that’s where I’ll settle down.
010
Mike Hearn @mikehearn.bsky.social · 16/01/2025
I might be out of touch, but at best I'd say it's tolerated. Not offensive enough for Apple users to disable en masse, basically, but not winning the hearts and minds of users (and also not making sales). 9to5mac.com/2025/01/10/a...
9to5mac.com
Apple Intelligence isn't helping Apple boost iPhone sales
Apple last year introduced Apple Intelligence, its own set of AI tools. Since they are all processed on-device, Apple Intelligence...
000