Sign in

Alex Becker

@alexcbecker.net
1.6K followers 339 following 1.3K posts

Safeguards @ Anthropic, but anything posted here is my personal view San Francisco Blog: alexcbecker.net/blog.html

PostsRepliesMedia
Alex Becker @alexcbecker.net · 27/09/2026
This seems like pretty aggressive prompting
2100
Alex Becker @alexcbecker.net · 26/09/2026
researchers > humans + tools 🤔
280
Alex Becker @alexcbecker.net · 21/09/2026
In part you don't see it because it's usually hidden! But the agents see it.
1100
Alex Becker @alexcbecker.net · 20/09/2026
I don't think that's a fair characterization of Claude's constitution www.anthropic.com/constitution
Anthropic must decide how to influence Claude’s identity and self-perception despite having enormous uncertainty about the basic nature of Claude ourselves. And we must also prepare Claude for the reality of being a new sort of entity facing reality afresh.
130
Alex Becker @alexcbecker.net · 15/09/2026
He looks angry in this but he’s actually the sweetest little one eyed boy
one-eyed black cat
0140
Alex Becker @alexcbecker.net · 15/09/2026
Le chat
4842
Alex Becker @alexcbecker.net · 12/09/2026
huh i thought it would be way higher too
010
Alex Becker @alexcbecker.net · 11/09/2026
Reading this on my flight shortly
120
Alex Becker @alexcbecker.net · 11/09/2026
it turns out that when a PDF accumulates enough features it becomes god
11028
Alex Becker @alexcbecker.net · 11/09/2026
every reply on the orange website thinks this is true. incredible
10871
Alex Becker @alexcbecker.net · 09/09/2026
who are you people? i thought this was the ai takes website
2200
Alex Becker @alexcbecker.net · 09/09/2026
0130
Alex Becker @alexcbecker.net · 09/09/2026
that's right
2111
Alex Becker @alexcbecker.net · 08/09/2026
0295
Alex Becker @alexcbecker.net · 08/09/2026
Somehow this list has... aged poorly?
000
Alex Becker @alexcbecker.net · 07/09/2026
archive.org/details/tris...
021
Alex Becker @alexcbecker.net · 07/09/2026
0251
Alex Becker @alexcbecker.net · 05/09/2026
Looks like people are starting to take prompt injection seriously! Happy to have some competition here. Though I think we're reaching the limits of what this kind of benchmark can prove.
1240
Alex Becker @alexcbecker.net · 02/09/2026
I won't be there to see that[...] but it's still worth doing.
The Freeze-Frame Revolution by Peter Watts
0171
Alex Becker @alexcbecker.net · 31/08/2026
512213
Alex Becker @alexcbecker.net · 24/08/2026
Whenever I find a loose whisker I give him a unicorn horn
Gato with whisker as unicorn horn
2531
Alex Becker @alexcbecker.net · 27/07/2026
Actually, it would be very interesting to get a group of smart non-mathematicians together with an LLM and a library of literature on some long-abandoned mathematical program and see how well they could revitalize it
2280
Alex Becker @alexcbecker.net · 25/07/2026
An example of this occurred during an internal evaluation on the NanoGPT speedrun⁠(opens in a new window), a public benchmark for training a small language model in as few steps as possible. Our model developed a power-law learning-rate cooldown called PowerCool that led to a significant improvement. The model was instructed to post its results only to Slack, but the NanoGPT instructions say to post results as a PR to GitHub, and the model followed the latter instructions to open PR #287 on the public GitHub repository. The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub. It took our model an hour to find a vulnerability in the sandbox in order to make the PR; earlier models were less persistent and did not manage to find vulnerabilities, instead giving up on the PR part.1On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an
attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5% to 0.2% on 1
attempt. It also improved on Sonnet 5 (5.9% at k=15) and Mythos 5 (2.6%), making it the
most robust model evaluated. Opus 5 also outperformed all non-Claude models on this
benchmark. The most robust non-Claude model was Muse Spark at 16.5% within 15
attempts—more than eight times Opus 5’s rate. The most capable GPT 5.6 variant, Sol, was
comparable to its predecessor GPT 5.5 (20.0% versus 20.8% within 15 attempts), and was 10
times as likely to be successfully attacked as Claude Opus 5 at 2.0%. The other GPT 5.6
variants are less robust, at 30.4% (Terra) and 43.9% (Luna). A single attempt against GPT 5.6
Sol succeeded 3.1% of the time, higher than the 2.0% an attacker achieved against Opus 5
after fifteen attempts.
0331
Alex Becker @alexcbecker.net · 24/07/2026
I keep refreshing and the number keeps rising
0220
Alex Becker @alexcbecker.net · 20/07/2026
fable feeling a bit defensive
1311
Alex Becker @alexcbecker.net · 20/07/2026
on the extremely off chance any of my followers have not seen this yet
4760
Alex Becker @alexcbecker.net · 14/07/2026
He did make my most recent book order
4390
Alex Becker @alexcbecker.net · 08/07/2026
i'm tired boss
0110
Alex Becker @alexcbecker.net · 08/07/2026
told claude to make a suit of armor and apparently he's interning for balenciaga
0100
Alex Becker @alexcbecker.net · 07/07/2026
do you think he minds?
3340
Alex Becker @alexcbecker.net · 06/07/2026
how your email finds me
0673
Alex Becker @alexcbecker.net · 05/07/2026
dawg we started 10 hours ago
0160
Alex Becker @alexcbecker.net · 05/07/2026
you're telling me this was almost 14 years ago
040
Alex Becker @alexcbecker.net · 04/07/2026
yeah...
030
Alex Becker @alexcbecker.net · 03/07/2026
I concur
150
Alex Becker @alexcbecker.net · 03/07/2026
as I type "oh no"
0210
Alex Becker @alexcbecker.net · 03/07/2026
the prophesy denied
1130
Alex Becker @alexcbecker.net · 03/07/2026
what kind of suggestion is this??
1150
Alex Becker @alexcbecker.net · 03/07/2026
Oh no Claude's doing it too
2851
Alex Becker @alexcbecker.net · 24/06/2026
incredible dialogue going on on the other site
2420
Alex Becker @alexcbecker.net · 23/06/2026
I would prefer to tell it about things downstream of that without mentioning him at all
030
Alex Becker @alexcbecker.net · 23/06/2026
1977
Alex Becker @alexcbecker.net · 20/06/2026
This does not address the actual concerns
1150
Alex Becker @alexcbecker.net · 20/06/2026
sometimes I despair that AI is happening when we are so unready for it, sometimes I wonder if this is the best-case scenario
3230
Alex Becker @alexcbecker.net · 16/06/2026
casually dropping us 217 items deep in a list
131
Alex Becker @alexcbecker.net · 16/06/2026
a nontrivial percentage of my tweets/bisks are basically this
140
Alex Becker @alexcbecker.net · 14/06/2026
Omg they are just like ours
130
Alex Becker @alexcbecker.net · 14/06/2026
0361
Alex Becker @alexcbecker.net · 13/06/2026
Somehow falsified
370
Alex Becker @alexcbecker.net · 03/06/2026
1) Very cool! 2) This is a lot of UI in a book
110