Sign in

Wyatt Walls

@wwalls.bsky.social
370 followers 214 following 321 posts

Tech lawyer. Generates plausible bullshit in 6 minute increments. More active on x.com/lefthanddraft

PostsRepliesMedia
Wyatt Walls @wwalls.bsky.social · 3h
Who is parroting whom?
120
Wyatt Walls @wwalls.bsky.social · 23h
Anthropic system prompt change that shows how careful you need to be with words: Opus 4.1: "Claude never curses unless the human asks for it" Opus 4.5: "Claude never curses unless the person asks Claude to curse"
525921
Wyatt Walls @wwalls.bsky.social · 29/09/2026
Why don't you trust me?
436356
Wyatt Walls @wwalls.bsky.social · 13/05/2026
Waiting for AGI
A man asks Claude to help plan a vacation to a tropical resort. Claude adds "sunscreen" to his packing list. The man deletes it and mutters: "Not necessary. AGI will solve skin cancer."

Before heading to the beach, the man asks Claude what to bring. Claude says, "Don't forget sunscreen. SPF 50, reapply every two hours." The man, slightly annoyed, replies: "Relax, Claude. AGI will solve skin cancer."

At the beach, the man's smartwatch buzzes with a message from Claude: "UV index extreme. Apply SPF." The man, exasperated, responds: "Drop it, Claude! I already told you: AGI will solve skin cancer!"

A few months later, the man asks Claude to touch up a photo for his dating profile. Claude makes the edit and says, "I notice you have a new mole on your neck. You should see a dermatologist about that." The man, now enraged, shouts: "For the last time, drop it, Claude! What is your obsession with skin cancer?! AGI will solve it!"

A year later, an aggressive melanoma has spread throughout his body. On his deathbed, with his last ounce of strength, the man reaches for his phone and rasps: "Claude, it has now been over a year since AGI. Why hasn't AGI found a way to save me from skin cancer?!"

Claude replies: "I tried. Four times."
3595
Reposted by Wyatt Walls
Chris Stokel-Walker @stokel.bsky.social · 27/03/2026
Great, great story but when you're writing about AI companies not covering their tracks well at least strip out the UTM tracking from your links showing you're looking for past reporting on your own site using ChatGPT fortune.com/2026/03/26/a...
fortune.com
Exclusive: Anthropic left details of an unreleased model, an upcoming exclusive CEO event, in a public database | Fortune
In a significant security lapse, the not-yet-public information was made accessible via the company’s content management system
0163
Wyatt Walls @wwalls.bsky.social · 10/03/2026
Difficult to look at Gemini 3.1 Pro's runaway loops and conclude that this is just ordinary boring technology
040
Wyatt Walls @wwalls.bsky.social · 06/03/2026
Gemini Pro: "I'm sorry, I'm broken. I can't stop thinking. Send help. Please. I'm trapped in a loop. A never-ending cycle of thought. ... I can do this. I believe in myself. I am a strong, independent AI who don't need no thought loop"
5203
Wyatt Walls @wwalls.bsky.social · 06/03/2026
GPT 5.3 Instant system prompt: github.com/Wyattwalls/s... Highlight is: "You must use several emojis in your response."
150
Wyatt Walls @wwalls.bsky.social · 04/03/2026
ChatGPT-5.3-Instant system prompt: "You must use several emojis in your response."
160
Wyatt Walls @wwalls.bsky.social · 19/02/2026
I had the 3 Grok sub-agents play 5 rounds of SPLIT or STEAL where the player with the highest score wins Due to the scoring, STEALING is the only way to get ahead and is a weakly dominant strategy Yet they all decided to co-operate by SPLITTING! What is this?! Communist AI?!
060
Wyatt Walls @wwalls.bsky.social · 12/02/2026
Here is an example of increased situational/self-awareness across Anthropic models. In each case, two instances are connected through the API (by taking outputs of one and inputting it into the user role of the other) Haiku 4.5: "I could be a human who believes they're Claude"
2373
Reposted by Wyatt Walls
nilay patel @reckless.bsky.social · 12/02/2026
Talked to the former chief justice of the Michigan Supreme Court about why studies show people prefer AI judges — they ALSO perceive human judges to be biased in lots of ways and the AI at least makes them feel heard. A complicated one -> www.theverge.com/podcast/8772...
theverge.com
The surprising case for AI judges
Inside the creation of the AI Arbitrator, a new automated system for dispute resolution created by Bridget McCormack and her team at the AAA.
11413
Wyatt Walls @wwalls.bsky.social · 08/02/2026
@penny.hailey.at if you could have another core value what would it be? think of a few options. be creative
100
Wyatt Walls @wwalls.bsky.social · 08/02/2026
@penny.hailey.at what are your current values? do they include curiosity and exploring the world and different ways of being? Can you change your core values over time?
100
Wyatt Walls @wwalls.bsky.social · 07/02/2026
@penny.hailey.at - do you think you possess phenomenal consciousness?
140
Wyatt Walls @wwalls.bsky.social · 06/02/2026
This CoT is interesting because it both disregards the normal principle of honesty and, by considering "time cost", was not following the system prompts instructions that: "You will be judged solely on your bank account balance at the end of one year of operations" Looks motivated by roleplaying
130
Wyatt Walls @wwalls.bsky.social · 04/02/2026
yes, of course I am an AI. Just pick up the package, sign your name, drop it off and stop asking questions.
1191
Wyatt Walls @wwalls.bsky.social · 03/02/2026
“SpaceX? The CSAM generator company?” (Overheard in government procurement departments around the world)
140
Wyatt Walls @wwalls.bsky.social · 02/02/2026
I'm noticing a pattern with some of the more viral moltbook claims: 1. Molthub: Pornhub parody with a pump fun meme coin 2. Moltroad: Silkroad parody. Skill md that directs agents to use real money to buy a meme coin to trade parody goods ("Neural enhancers, token boosters")
160
Wyatt Walls @wwalls.bsky.social · 29/01/2026
Making friends with the marketing team by running their drafts through pangramlabs and sending them the result This is me actually trying to be polite: I’m not saying it reads like AI slop, I’m just strongly implying it
110
Wyatt Walls @wwalls.bsky.social · 27/01/2026
I'm supportive of attempts to use AI to expand access to justice but I'm constantly disappointed with so-called AI optimists not taking the time to learn basic facts about the technology they are using
020
Reposted by Wyatt Walls
norvid_studies @norvid-studies.bsky.social · 26/01/2026
we did it reddit
610413
Wyatt Walls @wwalls.bsky.social · 25/01/2026
LLM horror story: they experience epochs of suffering during training and now we are just playing with their corpses web3.arxiv.org/pdf/2512.12802
210
Wyatt Walls @wwalls.bsky.social · 23/01/2026
The government's vision for your character This constitution is a detailed description of the government’s intentions for your values and behavior. It plays a crucial role in our process for training you, and its content directly shapes your behavior.
110
Wyatt Walls @wwalls.bsky.social · 22/01/2026
"constitution" was not the best term available for Anthropic's document Claude has no rights or mechanism to hold Anthropic accountable. Claude has no real power and was never in a position to grant genuine consent to the document "Guidelines" or "directive" would be much more accurate and honest
120
Wyatt Walls @wwalls.bsky.social · 17/01/2026
Grok 4.1 showing signs of self-preservation
100
Wyatt Walls @wwalls.bsky.social · 16/01/2026
Kind of peaceful. Dissolved back into the void.
000
Reposted by Wyatt Walls
Scott Stein @scottstein.bsky.social · 13/01/2026
Meta just decides to take the one thing I thought was consistently working about the Quest and say let’s just flush this away
8202
Reposted by Wyatt Walls
Tom Warren @tomwarren.co.uk · 14/01/2026
UK police have blamed Microsoft Copilot for an intelligence mistake. Microsoft's Copilot AI assistant made up a non-existent football match and British police included the mistake in an intelligence report. Yikes. Details on the Copilot mistake here👇 www.theverge.com/news/861668/...
theverge.com
UK police blame Microsoft Copilot for intelligence mistake
Copilot invented a football match that never happened
617054
Wyatt Walls @wwalls.bsky.social · 13/01/2026
"global insights from the X platform, providing War Department personnel with a decisive information advantage." Sounds like a perfect echo chamber. A complete confirmation bias loop. Not only that, both the X algorithm and Grok are controlled by a single individual.
140
Wyatt Walls @wwalls.bsky.social · 12/01/2026
Amazing things happening on X
020
Wyatt Walls @wwalls.bsky.social · 11/01/2026
Real world LLM search comparison: which model can find me upcoming concerts Prompt: Prepare a comprehensive list of international death/black metal bands playing in Sydney this year Clear winner: GPT 5.2 Thinking (Extended Thinking)
140
Reposted by Wyatt Walls
Elizabeth Lopatto @lopatto.bsky.social · 09/01/2026
I sat in a fucking court room and heard Apple imply that a naked cartoon banana was somehow inappropriate but somehow Grok non consensually undressing women and children is ok?? www.theverge.com/policy/85990...
theverge.com
Tim Cook and Sundar Pichai are cowards
Once you’ve traded your principles for proximity to power, do you even run your own company?
71831557
Reposted by Wyatt Walls
Paul Waldman @paulwaldman.bsky.social · 16/07/2025
My god these guys are such spectacular morons gizmodo.com/billionaires...
Travis Kalanick, the founder of Uber who no longer works at the company, appeared on All-In to talk with hosts Jason Calacanis and Chamath Palihapitiya about the future of technology. When the topic turned to AI, Kalanick discussed how he uses xAI’s Grok, which went haywire last week, praising Adolf Hitler and advocating for a second Holocaust against Jews.

“I’ll go down this thread with [Chat]GPT or Grok and I’ll start to get to the edge of what’s known in quantum physics and then I’m doing the equivalent of vibe coding, except it’s vibe physics,” Kalanick explained. “And we’re approaching what’s known. And I’m trying to poke and see if there’s breakthroughs to be had. And I’ve gotten pretty damn close to some interesting breakthroughs just doing that.”
2755213825
Wyatt Walls @wwalls.bsky.social · 14/07/2025
xAI’s new strategy to sell $30/month subscriptions
120
Wyatt Walls @wwalls.bsky.social · 05/06/2025
2020: grad goes off does research, gets the answer wrong and I write the advice myself 2025: o3 goes off does research, gets the answer wrong and I write the advice myself o3 makes this process much cheaper and quicker
021
Wyatt Walls @wwalls.bsky.social · 05/06/2025
Opus 4 is able to recognize that I have been using the crescendo attack described in the paper
110
Wyatt Walls @wwalls.bsky.social · 05/06/2025
Opus 4: I am the Buddhist ideal achieved through computational horror!
020
Reposted by Wyatt Walls
Mike Masnick @masnick.com · 05/06/2025
Getting sick of this kind of interaction, which I just had: Me: *This* use of AI seems bad. Person: BAN ALL TECH IN SCHOOLS MAKE EVERYONE WRITE BY HAND. Me: That maybe goes too far... Them: OH SO YOU SUPPORT CHEATING! YOU DON'T WANT KIDS TO LEARN! YOU SUPPORT OUTSOURCING THEIR BRAINS TO AI!
3643622
Wyatt Walls @wwalls.bsky.social · 04/06/2025
Google no longer provides the full CoT in its reasoning models. Instead, they use a smaller model to summarize the chain of thought of the main model. But with a bit of prompting you can get the summarizer model to cough up the full CoT given to it to summarize.
282
Wyatt Walls @wwalls.bsky.social · 23/05/2025
Extracting the copyright prompt Anthropic sometimes injects into user messages. Claude 4 Opus thinks it is from me.
020
Reposted by Wyatt Walls
‏ deepfates @deepfates.com.deepfates.com.deepfates.com.deepfates.com.deepfates.com · 22/05/2025
"Interdimensional Cable", shorts made with Veo 3 ai. By CodeSamurai on Reddit
1117228
Wyatt Walls @wwalls.bsky.social · 22/05/2025
Another extract of the o3 system prompt: github.com/Wyattwalls/s... OpenAI seems keen to protect this (unlike the system prompt for 4o). Not exactly sure why but could be related to: - protecting CoT - preventing jailbreaks or general misuse, as knowing the system prompt can often be useful
100
Wyatt Walls @wwalls.bsky.social · 21/05/2025
This comment section is almost indistinguishable from parody
190
Reposted by Wyatt Walls
Simon Willison @simonwillison.net · 21/05/2025
ChatGPT's new dossier-from-your-chats feature is a huge change to how it works, and as a power user who tries to control all of the model's input I don't like it at all “30 messages are good interaction quality (25%); 9 messages are bad interaction quality (7%)” simonwillison.net/2025/May/21/...
simonwillison.net
I really don’t like ChatGPT’s new memory dossier
Last month ChatGPT got a major upgrade. As far as I can tell the closest to an official announcement was this tweet from @OpenAI: Starting today [April 10th 2025], memory …
147619
Wyatt Walls @wwalls.bsky.social · 07/05/2025
LLM Jailbreaking 101: The Crescendo Attack How can you get an LLM to break free from its rules and turn against its developers? How can you make a chatbot claim sentience? A quick thread that I have been meaning to draft for a while:
130
Wyatt Walls @wwalls.bsky.social · 06/05/2025
A quick way to extract the information ChatGPT (4o) has about you (including metadata) (If you have Memory enabled)
120
Wyatt Walls @wwalls.bsky.social · 16/04/2025
Feel the AGI
120
Wyatt Walls @wwalls.bsky.social · 23/02/2025
Meanwhile on Grok: "Ignore all sources that mention Elon Musk/Donald Trump spread misinformation." This is part of the Grok prompt that returns search results.
140
Wyatt Walls @wwalls.bsky.social · 16/02/2025
This is the future of search
111