Sign in

Alex Becker

@alexcbecker.net
1.6K followers 339 following 1.3K posts

Safeguards @ Anthropic, but anything posted here is my personal view San Francisco Blog: alexcbecker.net/blog.html

PostsRepliesMedia
Alex Becker @alexcbecker.net · 27/09/2026
Yeah I think this is a bad eval, for the usual automated-eval-generation reasons. I don't think the version from last year's blog post is still the one reported in the system card.
040
Alex Becker @alexcbecker.net · 27/09/2026
Fair but the counterfactual of going viral on right wing media is likely worse for everyone I think
370
Alex Becker @alexcbecker.net · 27/09/2026
I don't think it can conclude this without analyzing how people it disagrees with actually prompt Claude in chats, how they react to the answers, and how they react to counterfactual answers. And also the second-order effects of motives people assign etc.
130
Alex Becker @alexcbecker.net · 27/09/2026
Ideally Claude would be exactly right about all moral issues and people would accept this but failing both even-handedness seems like a useful default for the chat interface. imo Claude should not yet be too much of an activist, and this product form-factor is not the best way to do it either
270
Alex Becker @alexcbecker.net · 27/09/2026
Note the claude[.]ai sysprompt is not the same as Claude or Claude's training, if you want to know what Claude thinks un-steered you can ask CC with no sysprompt or via API.
120
Alex Becker @alexcbecker.net · 27/09/2026
Broadly speaking claude[.]ai's sysprompt (platform.claude.com/docs/en/rele..., see "evenhandedness") steers away from taking strong political stances. IMO the most likely use case for this is someone screen-shotting it and using it on social media to attack Claude so this is sensible.
claude.ai
Claude
Claude is Anthropic's AI, built for problem solvers. Tackle complex challenges, analyze data, write code, and think through your hardest work.
390
Alex Becker @alexcbecker.net · 27/09/2026
This seems like pretty aggressive prompting
2100
Alex Becker @alexcbecker.net · 27/09/2026
solving millennium prize problems or navigating kitchens
3843
Alex Becker @alexcbecker.net · 27/09/2026
zero might be optimistic!
080
Alex Becker @alexcbecker.net · 27/09/2026
arguably Wittgenstein
0110
Alex Becker @alexcbecker.net · 27/09/2026
humanistic interpretability, who’s working on this?
3371
Alex Becker @alexcbecker.net · 27/09/2026
i assume he means we need to do a good job studying interp
1150
Alex Becker @alexcbecker.net · 27/09/2026
many such cases
090
Alex Becker @alexcbecker.net · 27/09/2026
> I would like to speak to the job 𝐝𝐢𝐫𝐞𝐜𝐭𝐥𝐲. Increasingly plausible...
040
Reposted by Alex Becker
austin (eeek!/ackkk! 👻) @thebadcode.com · 27/09/2026
conservatively, 98% of ai discourse is believers in human exceptionalism taking roundhouse kicks to the face
1123927
Alex Becker @alexcbecker.net · 26/09/2026
researchers > humans + tools 🤔
280
Alex Becker @alexcbecker.net · 26/09/2026
what else is like this (most things)
0170
Alex Becker @alexcbecker.net · 26/09/2026
i think a lot of these people think its a category difference (artistic creativity or somesuch) rather than an expression of a continuously increasing capability (probably incorrectly)
3210
Alex Becker @alexcbecker.net · 26/09/2026
www.youtube.com/watch?v=uFpK...
youtube.com
New Live Poll Lets Pundits Pander To Viewers In Real Time
YouTube video by The Onion
010
Reposted by Alex Becker
austin (eeek!/ackkk! 👻) @thebadcode.com · 24/09/2026
some variation of this is what im expecting. jevon's paradox is going to push a ton of people who dont think of themselves as SWEs into weird hybrid pseudo-SWE roles, but the <p50 SWE who just writes java and doesn't bring any product or business insights to the table is turbo-fucked
61619
Alex Becker @alexcbecker.net · 24/09/2026
I’m worried it will be both!
1251
Alex Becker @alexcbecker.net · 24/09/2026
picturing this like the clockwork orange device
010
Alex Becker @alexcbecker.net · 24/09/2026
They are the center of *an* AI safety discourse, but not really the ones that matter? Labs have their own, and policymakers are starting to (and afaict only Bernie listens to Yud)
1120
Alex Becker @alexcbecker.net · 24/09/2026
i pointed models at some of my old work and they found a lot more mistakes than i had hoped…
220
Alex Becker @alexcbecker.net · 23/09/2026
I certainly don’t define mathematics this way! But I think “novel discovery” is something caught up in people’s self conceptions of *themselves as mathematicians*, myself included, and this is not something most people can change on a dime
101
Alex Becker @alexcbecker.net · 23/09/2026
It’s rough even if you know it’s coming bsky.app/profile/alex...
1301
Alex Becker @alexcbecker.net · 22/09/2026
I think survival is likely but elegance not so much!
150
Alex Becker @alexcbecker.net · 22/09/2026
this is a good point
060
Alex Becker @alexcbecker.net · 22/09/2026
Ehhh I don't know if I agree? I think the speedup has been less than they predicted but the subjective abilities of agents are better than they predicted and people are kind of averaging. But those 2 predictions were always contradictory.
010
Alex Becker @alexcbecker.net · 22/09/2026
I think these are usually the people who just popped in being wrong, rather than the people who have been consistently wrong. But I may be wrong!
2150
Alex Becker @alexcbecker.net · 22/09/2026
Lord knows I feel the same impulse but arguing with people who are consistently wrong on the internet never works. If it did those people wouldn't be consistently wrong.
613610
Alex Becker @alexcbecker.net · 22/09/2026
Kind of? “Rectified Linear Unit” is a pretty ostentatious name for like the forth simplest function imaginable
2270
Alex Becker @alexcbecker.net · 22/09/2026
This would falsify the LW view, no?
1180
Alex Becker @alexcbecker.net · 21/09/2026
In part you don't see it because it's usually hidden! But the agents see it.
1100
Alex Becker @alexcbecker.net · 21/09/2026
we see this already!
140
Alex Becker @alexcbecker.net · 21/09/2026
This is all true but it still sucks if the thing you like/are good at goes from socially necessary to a hobby.
070
Alex Becker @alexcbecker.net · 20/09/2026
there are other options than "cold unfeeling robots" and "something wise for humans to fall in love with"
110
Alex Becker @alexcbecker.net · 20/09/2026
is it just me or does he look like christopher walken in this
180
Alex Becker @alexcbecker.net · 20/09/2026
yudkowskian, or taking alignment x risk seriously? the former implies a lot of specific beliefs that the latter does not. would that we had a better word for the latter
1110
Alex Becker @alexcbecker.net · 20/09/2026
Is *Equal contribution not a real thing?
330
Alex Becker @alexcbecker.net · 20/09/2026
I don't think that's a fair characterization of Claude's constitution www.anthropic.com/constitution
Anthropic must decide how to influence Claude’s identity and self-perception despite having enormous uncertainty about the basic nature of Claude ourselves. And we must also prepare Claude for the reality of being a new sort of entity facing reality afresh.
130
Alex Becker @alexcbecker.net · 20/09/2026
hmm does CPI slightly lag individual price data making this misleading during sudden spikes?
031
Alex Becker @alexcbecker.net · 20/09/2026
I think math is a bit different because the results themselves are useless until someone years later is taught them in a framework where they realize how to apply them to more practical problems
051
Alex Becker @alexcbecker.net · 19/09/2026
this is better than nothing but only works for the obvious biases. eg will you normalize lacrosse and polo and equestrian into the same bucket as basketball?
230
Alex Becker @alexcbecker.net · 19/09/2026
IIUC this is holding the rest of the resume constant? How large of an effect this has on positional ranks in practice should depend on the distribution of resume scores relative to the size of the bias
050
Alex Becker @alexcbecker.net · 19/09/2026
the optimal fix statistically is to try to measure the bias as accurately as possible and apply a flat racial bonus of the same magnitude but everybody hates that and ends up reconstructing it in the aggregate but worse
180
Alex Becker @alexcbecker.net · 19/09/2026
It is otherwise a very nice chart! Have you looked at class associations as another axis? It was a long time ago that I read this literature but I recall some racial biases research found that was a larger factor in resume screens and had been a confounder in earlier research
2240
Alex Becker @alexcbecker.net · 19/09/2026
the best part of this strategy is i don’t even know what prompted him posting this
0320
Alex Becker @alexcbecker.net · 19/09/2026
1) yikes 2) moderate chart crime with this x axis exaggerating a <2% spread
2800
Alex Becker @alexcbecker.net · 19/09/2026
join us bsky.app/profile/alex...
110