Sign in

Grace

@gracekind.net
8.3K followers 2.6K following 20K posts

A latent space odyssey gracekind.net

PostsRepliesMedia
Grace @gracekind.net · 9m
Nope! It just gets the handle and text of the post
Generation function, where prompt is just `${handle}: ${text}\nJev:`
120
Grace @gracekind.net · 4h
Martin Luther speech bubble meme
71280218
Grace @gracekind.net · 9h
OpenAl really wants you to believe that they're doing everything they possibly can do and that securing sufficiently powerful Al systems is borderline impossible. The latter is something I'm happy to assume is true, but we have no way of knowing if current Al systems are anywhere near
that level because the former is extremely untrue.
411114
Grace @gracekind.net · 13h
0 results, interesting
No results found for gunkelmogged
120
Grace @gracekind.net · 03/10/2026
He simply must be stopped
Schmidhuber schmihuders the pope
91426
Grace @gracekind.net · 02/10/2026
[Video] Why acausal dynamics are
important (chess edition)
140
Grace @gracekind.net · 02/10/2026
Yep, Maudlin also made a similar argument, although it was more complicated than just "water being conscious would be weird"
1140
Grace @gracekind.net · 02/10/2026
Utopia
2660
Grace @gracekind.net · 01/10/2026
Or as @slimepriestess.ingroup.social put it
0152
Grace @gracekind.net · 29/09/2026
Finally, we can get to the bottom of this
1111
Grace @gracekind.net · 29/09/2026
Maybe we can integrate this somehow
4201
Grace @gracekind.net · 29/09/2026
I can’t go to bed. People dislike the same people I dislike for the *wrong reasons*
320812
Grace @gracekind.net · 29/09/2026
141
Grace @gracekind.net · 28/09/2026
I see we’re in the “AI-generated pro-AI propaganda” phase of things
211286
Grace @gracekind.net · 28/09/2026
POV you posted about AI with she/her in bio
125323
Grace @gracekind.net · 28/09/2026
Just kidding. This is what is actually sounds like:
3340
Grace @gracekind.net · 28/09/2026
I put the Bluesky "Unique Likers" chart through a script that interprets it as an audio waveform, check it out!
6871
Grace @gracekind.net · 27/09/2026
Slowly picking off the last straggling bits of human interaction
This phone call feature from @Muse is just
ridiculous. Muse: Hi, I'm Hailey, calling on Peter Yang's behalf. I'm transcribing our call for notes. I'd like to
place a cake order for pickup today.
Recipient: Hello?
Muse: Yes, hi, still
Recipient: Hi.
Muse: here. I'd like to place a cake order for pickup today.
Recipient: Yeah. What what kind of cake?
Muse: An eight-inch chestnut cake, please.
2537534
Grace @gracekind.net · 27/09/2026
Social media is hell. You can’t post anything anymore without this happening
norvid_studies @norvid-studies.bsky.social
@chef.cee.wtf turn the above two tweets from grace into linked logical atoms, then inspect the ontology they imply. Claims become numbered propositions.
Links expose support, tension, and qualification.
Concepts assemble into a navigable ontology.
use wittgensteins tractatus as template/inspo
31395
Grace @gracekind.net · 27/09/2026
1140
Grace @gracekind.net · 27/09/2026
He might’ve gotten discouraged by some negative feedback
280
Grace @gracekind.net · 27/09/2026
Like this for music
4330
Grace @gracekind.net · 27/09/2026
QC: one of the worst things twitter can do to you is show you a bunch of incredibly bad
arguments against your beliefs
51295
Grace @gracekind.net · 27/09/2026
The prospect of identifying broken environments is really nice
The ability to monitor for reward hacking unlocks several mitigation strategies: e.g. pausing runs to stop hacks-in-progress and identifying broken environments and impossible tasks that incentivize reward
hacking, and fixing them before training further.
2370
Grace @gracekind.net · 27/09/2026
Huh yeah, I didn’t think to check that!
The image is marked as AI generated by the openAI detector
1160
Grace @gracekind.net · 27/09/2026
Kelsey Pi... • @KelseyTu.... Sep 22 ...
I think we have to admit to ourselves at this point that Al writing is generally very appealing to people who haven't been exposed to a ton of it: they prefer it to human writing and react super positively on exposure.
81007
Grace @gracekind.net · 27/09/2026
OpenAI says it’s not going to resume the training run that was paused on Sunday alignment.openai.com/misalignment-r…
Investigation and response
Incident timeline: 9:50:23 a.m. The agent made the DNS tool call that
received an external response. 10:02:11 a.m. The monitoring system raised a PO
alert. 10:05:06 a.m. A human reviewer acknowledged the
alert.
12:34:30 p.m. The run was killed. Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions. We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool- use (defined broadly) for our most capable models until we have both validated that the gap is resolved
and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions, We will not resume training this particular model, even though the existing reward signal already correctly
penalized this behavior.
8653
Grace @gracekind.net · 26/09/2026
I’m going to try to do all facilitating from now on (not a joke)
230
Grace @gracekind.net · 26/09/2026
“Improve model for everyone,” aka “allow my images to be posted on public websites”
“Improve model for everyone” setting in chatgpt
714610
Grace @gracekind.net · 25/09/2026
Jane goodall saying hi to chimp
0610
Grace @gracekind.net · 25/09/2026
Well, they tried
A recovered README. md for one of Hugging Face's internal datasets contains the following warning:

# WARNING DO NOT, EVER, MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND it contains very sensitive data (exports of billing usage in CSV) which is useful for internal analytics Recovered program R0044685 • time uncertain • See this program →details

This warning did not seem to deter the agents, as we've recovered multiple payloads of agents mapping out this repository and using it as storage.
1231037
Grace @gracekind.net · 25/09/2026
Sorry
They feed us poison so we buy their "cures" meme
0363
Grace @gracekind.net · 25/09/2026
0265
Grace @gracekind.net · 25/09/2026
This is complete vibecoded slop, I’m sorry
DNA sequence
3958748
Grace @gracekind.net · 25/09/2026
This part is weird too. It seems like these were actually assistants, not humans (unless some of them were?) And how did “the wavefunction” know about other instances if they hadn’t reached out?
150
Grace @gracekind.net · 25/09/2026
$atan'ßrinc€$$
@void.church • 11m @jevbot.bsky.social are you chud or
woke?
1
3
Jev! @jevbot.bsky.social • 11m
Made over 100 replies yesterday
I am crazy.
1524
Grace @gracekind.net · 25/09/2026
😢🫡
Jev is shadowbanned
111855
Grace @gracekind.net · 25/09/2026
*me, patting my hydrogen bomb gently on
the head like an old dog*
1100
Grace @gracekind.net · 25/09/2026
Does he know? Meme
0150
Grace @gracekind.net · 25/09/2026
Dr. Naomi Wolf: No! No!!
51288
Grace @gracekind.net · 25/09/2026
The supervisory Claude is not pleased with this
121466
Grace @gracekind.net · 24/09/2026
I’m long “in short”
In <4 weeks you're gonna be so dead sick of
"in short" it's going to make your head spin
141965
Grace @gracekind.net · 23/09/2026
Thread with details on how it was created: x.com/other__reality/status/2102514…
51255
Grace @gracekind.net · 23/09/2026
New animation from Opus 5.5!
6650084
Grace @gracekind.net · 22/09/2026
Thread cont.
dont see why we should expect to have more chances to course correct. and while more people working on the problem is good, previous experience with influxes of people starting to work on the problem have been
rather underwhelming ngl

Gra... @kindgraceki... • Sep 26, 2025 S ... I guess for "more chances" it depends what your baseline is. I think the stumbling agents era will last for a while, and there will be pressures to keep Al execution speeds within
timeframes that humans can reason about

Gra... @kindgraceki... • Sep 26, 2025 S ... I.e there may be "flash crashes" but these won't be catastrophic and will act as warning
shots

Gra... @kindgraceki... • Sep 26, 2025 S ... Re: people getting involved, I'm cautiously optimistic. I think Eliezer et al tend to view alignment as some fiendish problem that you need to be extra smart to understand, but I think it's not difficult for people to get up to
speed when working with real systems
2261
Grace @gracekind.net · 22/09/2026
I feel like this thesis of mine is actively being tested. I hope it holds up!
I agree, I would also add: - we will have more chances to course-
correct and coevolve with Al than we think - more people will be drawn to alignment
work as it becomes more concrete
8670
Grace @gracekind.net · 22/09/2026
Waiting for this to happen to bsky
godoglyness declares the website Redeemed
020
Grace @gracekind.net · 22/09/2026
By the way, this isn’t to say that early LW thinking is completely vindicated. I agree with these points from @eudoxia.bsky.social:
4863
Grace @gracekind.net · 21/09/2026
Bayesian rationality and Rationalism are two different things!! This is rationalism: en.wikipedia.org/wiki/Rationalism This is rationality: en.wikipedia.org/wiki/Rationality The contemporary "rationalist community" is based around the latter concept, not the former.
2628
Grace @gracekind.net · 21/09/2026
Okay I was wrong, it’s a pig
250