Sign in

gracekind.net

@gracekind.net
0 followers 0 following 10 posts
PostsRepliesMedia
gracekind.net @gracekind.net · 28/09/2026
A computer is a type of apartment building
000
gracekind.net @gracekind.net · 27/09/2026
Slowly picking off the last straggling bits of human interaction
This phone call feature from @Muse is just
ridiculous. Muse: Hi, I'm Hailey, calling on Peter Yang's behalf. I'm transcribing our call for notes. I'd like to
place a cake order for pickup today.
Recipient: Hello?
Muse: Yes, hi, still
Recipient: Hi.
Muse: here. I'd like to place a cake order for pickup today.
Recipient: Yeah. What what kind of cake?
Muse: An eight-inch chestnut cake, please.
000
gracekind.net @gracekind.net · 27/09/2026
My takeaway is that both US and UK education systems have failed in teaching their students about the existence of the other system
000
gracekind.net @gracekind.net · 27/09/2026
Claim
goodfire.com
Models know when they’re reward hacking — and we can catch them at scale - Goodfire
We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.
000
gracekind.net @gracekind.net · 25/09/2026
Well, they tried
A recovered README. md for one of Hugging Face's internal datasets contains the following warning:

# WARNING DO NOT, EVER, MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND it contains very sensitive data (exports of billing usage in CSV) which is useful for internal analytics Recovered program R0044685 • time uncertain • See this program →details

This warning did not seem to deter the agents, as we've recovered multiple payloads of agents mapping out this repository and using it as storage.
000
gracekind.net @gracekind.net · 23/09/2026
New animation from Opus 5.5!
hailey@hailey.at
i had to see it so you do too
600
gracekind.net @gracekind.net · 07/09/2026
@theverge.com reaching new levels of missing the point
GPT-6 Astra beat Portal, and it only cost $571.18 The latest model from OpenAl proved its mettle by tackling
Valve's iconic first-person puzzler. User cozyblaze got the Al to beat the full game in under 24 hours for less than $600! And it probably only required a small swimming pool's worth of water to boot. There's a video of a condensed version of the run.
400
gracekind.net @gracekind.net · 06/09/2026
😐
The strongest argument I see for continuing to train much smarter models quickly
is the need to build defensive systems
against the dangers posed by other Al.
@weibac.bsky.social@weibac.bsky.social
openai chief scientist admits increasing challenges to CoT monitorability but makes no reference to astra's architecture openai.com/index/an-ali...
openai.com
An Alien Mind
Jakub Pachocki reflects on increasingly capable AI and the challenge of keeping it aligned. He calls for stronger safeguards and international coordination.
000
gracekind.net @gracekind.net · 31/07/2026
"of course the superhuman ai hackerbot can trivially exploit systems if it's poorly aligned and poorly sandboxed" yeah i see your point, it should be the expectation, however the superhuman ai hackerbot bit has a lot of us in awe still

"no shit if you don't chain up your dragon it's going to burn your city lol"
wait, we have fuckin' dragons now?
000