Sign in

Simon Willison

@simon.fedi.simonwillison.net.ap.brid.gy
10K followers 3 following 1.6K posts

Open source developer building tools to help journalists, archivists, librarians and others analyze, explore and publish their data. datasette.io […] [bridged from fedi.simonwillison.net/@simon on the fediverse by fed.brid.gy ]

PostsRepliesMedia
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 17/08/2026
My conclusions from the end if the post
Some observations #
The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I’m delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models—today it can run on a capable laptop.

The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That’s the catch with these dense (non-Mixture-of-Experts) models—they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard.

The most important thing about Qwen 3.8 27B is what it demonstrates. We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file.

The models at this size continue to get better at an impressive rate. We don’t need to spend half a million dollars on datacenter-class hardware just to run a competent model.
081
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 14/08/2026
The new Qwen 3.7 27B, running as a 17GB GGUF in LM Studio on my M5 Max MacBook Pro, just drew me the best pelican riding a bicycle I've seen from any model that runs on my laptop It did take nearly 21 minutes to generate, and used 22,276 reasoning […] [Original post on fedi.simonwillison.net]
The pelican has the right shaped beak. The red bicycle has the correct shape of frame. The pelican's wing reaches the handlebars. It has legs on both side of the bicycle. There is a pleasing set of clouds, birds, sun, grass and shadow on the image, plus motion lines behind but not in front of the bird.
293
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 11/06/2026
Since the JSON extras API is a little hard to explain without an example I also had Fable 5 and GPT-5.5 collaborate on this custom API explorer tool for trying out the new feature […] [Original post on fedi.simonwillison.net]
Screenshot of a web application titled "Datasette extras explorer". A URL input field contains https://latest.datasette.io/fixtures/facetable.json with a teal Explore button next to it. Below, a left panel labeled EXTRAS (30) lists checkboxes: all_columns - All columns in the table, regardless of _col/_nocol filtering; column_types - Column type assignments for this table; columns (checked) - Column names returned by this query; count - Total count of rows matching these filters; count_sql - SQL query used to calculate the total count; custom_table_templates - Custom template names considered for this table; database - Database name; database_color - Color assigned to the database. A right panel labeled RESPONSE shows GET /fixtures/fac… with Copy JSON and Copy URL buttons, then a dark JSON viewer showing 200 - 9.9 KB - 114ms and JSON: "ok": true, "next": null, "columns": (highlighted array) "pk", "created", "planet_int", "on_earth", "state", "_city_id", "_neighborhood", "tags", "complex_array", "distinct_some_null", "n", "rows": list of objects.
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 29/05/2026
Notes on Claude Opus 4.8, which Anthropic call "a modest but tangible improvement" - plus pelicans riding bicycles for each of the five different thinking efforts simonwillison.net/2026/May/28/claud…
SVGs of pelicans riding bicycles, in low, medium, high, xhigh and max. They do get progressively better. Only the max one has a correctly shaped bicycle frame.
3104
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 25/05/2026
Plugins can also affect the empty state of the new menu - the latest datasette-agent adds a form for kicking off a new agent conversation - live demo (if you sign in with GitHub) on agent.datasette.io
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 25/05/2026
Fun new feature in Datasette 1.0a30 - a "jump to" menu, also available by the "/" keyboard shortcut, which lets you type to jump to a database, table or canned query. It offers a hook so plugins can extend it with additional content too datasette.io/blog/2026/jump-menu
100
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 11/05/2026
New TIL: I figured out how to use my LLM CLI tool in a shebang line, which means you can write executable scripts in English, or hook up more complex scripts with a snippet of YAML template - til.simonwillison.net/llms/llm-sheb…
#!/usr/bin/env -S llm -x -f
Generate an SVG of a pelican riding a bicycle

But you can also incorporate tool calls:

#!/usr/bin/env -S llm -T llm_time -f
Write a haiku that mentions the exact current time

Or even execute YAML templates directly that define extra tools as Python functions:

#!/usr/bin/env -S llm -t
model: gpt-5.4-mini
system: |
  Use tools to run calculations
functions: |
  def add(a: int, b: int) -> int:
      return a + b
  def multiply(a: int, b: int) -> int:
      return a * b

Then:

./calc.sh 'what is 2344 * 5252 + 134' --td

Which outputs (thanks to that --td tools debug option):

Tool call: multiply({'a': 2344, 'b': 5252})
  12310688

Tool call: add({'a': 12310688, 'b': 134})
  12310822

2344 × 5252 + 134 = **12,310,822**
130
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 28/04/2026
It didn't quite manage to draw me a pelican riding a bicycle, but I still appreciated its era-appropriate response
I couldn't resist hitting the talkie demo with my classic test prompt:

Generate an SVG of a pelican riding a bicycle

It replied:

An SVG of a pelican riding a bicycle was generated in 1860. The pelican was represented as perched on the saddle, with its long bill pointed forward, and its legs astride the handles. The idea was suggested by the fact that pelicans are sometimes seen fishing on horseback, on the banks of the Rhine.
233
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 28/04/2026
Here's a uv one-liner that downloads and runs the MLX model against a local mp3 file uv run --with mlx-audio python -m mlx_audio.stt.generate \ --model mlx-community/VibeVoice-ASR-4bit \ --audio lenny.mp3 --output-path lenny \ --format json --verbose --max-tokens 32768
Screenshot of a macOS terminal running an mlx-audio speech-to-text command using the VibeVoice-ASR-4bit model on lenny.mp3, showing download progress, a warning that audio duration (99.8 min) exceeds the 59 min maximum so it's trimming, encoding/prefilling/generating progress bars, then a Transcription section with JSON segments of speakers discussing AI coding agents, followed by stats: Processing time 524.79 seconds, Prompt 26615 tokens at 50.718 tokens-per-sec, Generation 20248 tokens at 38.585 tokens-per-sec, Peak memory 30.44 GB.
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 27/04/2026
Couldn't resist capping that off with my all-time favorite quote from Matt Levine
And the investors wailed and gnashed their teeth but it’s true, that is what they agreed to, and they had no legal recourse. And OpenAI’s new CEO, and its nonprofit board, cut them a check for their capped return and said “bye” and went back to running OpenAI for the benefit of humanity. It turned out that a benign, carefully governed artificial superintelligence is really good for humanity, and OpenAI quickly solved all of humanity’s problems and ushered in an age of peace and abundance in which nobody wanted for anything or needed any Microsoft products. And capitalism came to an end.
170
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 24/04/2026
DeepSeek V4 just dropped - two models, Flash and Pro, both benchmarking well, decent pelicans and prices that put them both as the cheapest in their respective categories by a solid margin simonwillison.net/2026/Apr/24/deeps…
Flash: Excellent bicycle - good frame shape, nice chain, even has a reflector on the front wheel. Pelican has a mean looking expression but has its wings on the handlebars and feet on the pedals. Pouch is a little sharp.Pro: Another solid bicycle, albeit the spokes are a little jagged and the frame is compressed a bit. Pelican has gone a bit wrong - it has a VERY large body, only one wing, a weirdly hairy backside and generally loos like it was drown be a different artist from the bicycle.Table comparing AI model pricing with columns Model, Input ($/M), Output ($/M): DeepSeek V4 Flash $0.14 $0.28; GPT-5.4 Nano $0.20 $1.25; Gemini 3.1 Flash-Lite $0.25 $1.50; Gemini 3 Flash Preview $0.50 $3; GPT-5.4 Mini $0.75 $4.50; Claude 4.5 Haiku $1 $5; DeepSeek V4 Pro $1.74 $3.48; Gemini 3.1 Pro $2 $12; GPT-5.4 $2.50 $15; Claude Sonnet 4 / 4.5 $3 $15; Claude Opus 4.5 $5 $25; GPT-5.5 $5 $30.
092
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 20/04/2026
I upgraded my Claude token counter tool to compare different models and Opus 4.7 does appear to use 1.46x times the tokens for text and up to 3x the tokens for images - it's priced the same as Opus 4.6 on a per-token basis so this is actually a pretty […] [Original post on fedi.simonwillison.net]
Screenshot of a token comparison tool with an uploaded screenshot PNG image. Models to compare: claude-opus-4-7 (checked), claude-opus-4-6 (checked), claude-opus-4-5, claude-sonnet-4-6, claude-haiku-4-5. Note: "These models share the same tokenizer". Blue "Count Tokens" button. Results table — Model | Tokens | vs. lowest. claude-opus-4-7: 4,744 tokens, 3.01x (yellow badge). claude-opus-4-6: 1,578 tokens, 1.00x (green badge).
255
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 07/04/2026
I'm a big fan of the pelican GLM-5.1 drew me today, it even animated it! simonwillison.net/2026/Apr/7/glm-51
091
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 02/04/2026
Pelicans for Gemma 4 E2B, E4B, 26B-A4B and 31B - the first three generated on my laptop via LM Studio, the 31B was broken on my laptop so I ran it via the Gemini API instead simonwillison.net/2026/Apr/2/gemma-4
Two blue circles on a brown rectangle and a weird mess of orange blob and yellow triangle for the pelicanTwo black wheels joined by a sort of grey surfboard, the pelican is semicircles and a blue blob floating above itBicycle has the right pieces although the frame is wonky. Pelican is genuinely good, has a big triangle beak and a nice curved neck and is clearly a bird that is sitting on the bicycleMotion blur lines, a mostly great bicycle albeit missing the front part of the frame. Pelican is decent.
210
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 02/03/2026
Sent the February edition of my sponsors-only newsletter - a summary of my last month of blogging for people who want to pay for a shorter version I use Claude as a proofreader and fact checker, was delighted that it called me out on this Kākāpō […] [Original post on fedi.simonwillison.net]
5. "No new chicks for four years (due to a lack of fruiting rimu trees)" The phrasing "lack of fruiting rimu trees" is slightly imprecise. The issue isn't that rimu trees failed to fruit at all, but that there was no mass fruiting (masting) event, which is the specific trigger for kākāpō breeding. Consider "due to a lack of rimu masting" or "due to a lack of mass rimu fruiting."
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 27/02/2026
Built a fun prototype this morning of binary search using HTTP range requests, in this case to lookup characters in ~77MB of unicode data simonwillison.net/2026/Feb/27/unico…
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 21/02/2026
Fun bonus tip: regular Claude chat (not Claude Code) has the ability to clone repos from GitHub these days, which means you can ask it to checkout ANY public repo and answer questions about it or even use it as a starting point for an artifact!
Sonnet 4.6 
Extended
4 artifacts

User: Clone simonw/simonwillisonblog and tell me about the models and views

Thought: Synthesized comprehensive overview of models... 

Here's a comprehensive overview of the models and views in simonw/simonwillisonblog - a
Django-powered personal blog:
Models
blog app — the core
The blog is built around a shared BaseModel abstract class that all content types inherit from. It provides: created, slug, tags (M2M), metadata (JSON), search_document (Postgres full-text search vector), is_draft, card_image, and series (FK).
020
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 19/02/2026
Had to update my blog post after I saw Google's Jeff Dean had posted a video of an SVG animated pelican riding a bicycle, plus a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine
Screenshot of a blog post update. Text reads: "Update: In What happens if AI labs train for pelicans riding bicycles? last November I said:" followed by a blockquote: "If a model finally comes out that produces an excellent SVG of a pelican riding a bicycle you can bet I'm going to test it on all manner of creatures riding all sorts of transportation devices." Then: "Google's Gemini Lead Jeff Dean tweeted this video featuring an animated pelican riding a bicycle, plus a frog on a penny-farthing and a giraffe driving a tiny car and an ostrich on roller skates and a turtle kickflipping a skateboard and a dachshund driving a stretch limousine." Below are two side-by-side AI-generated SVG images labeled "Gemini 3 Pro" and "Gemini 3.1 Pro", both showing a dachshund driving a black stretch limousine. The Gemini 3 Pro version has a bright blue sky background, while the Gemini 3.1 Pro version has a sunset cityscape with a street lamp. At the bottom: "Prompt: Generate an animated 4:3 SVG of a dachshund driving a stretch limousine"
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 17/02/2026
New release of Rodney, my CLI tool for browser automation (designed for use by coding agents and with Showboat) - contributions from five people! simonwillison.net/2026/Feb/17/rodney


        Errors now use exit code 2, which means exit code 1 is just for for check failures. #15
        New rodney assert command for running JavaScript tests, exit code 1 if they fail. #19
        New directory-scoped sessions with --local/--global flags. #14
        New reload --hard and clear-cache commands. #17
        New rodney start --show option to make the browser window visible. Thanks, Antonio Cuni. #13
        New rodney connect PORT command to debug an already-running Chrome instance. Thanks, Peter Fraenkel. #12
        New RODNEY_HOME environment variable to support custom state directories. Thanks, Senko Rašić. #11
        New --insecure flag to ignore certificate errors. Thanks, Jakub Zgoliński. #10
        Windows support: avoid Setsid on Windows via build-tag helpers. Thanks, adm1neca. #18
        Tests now run on windows-latest and macos-latest in addition to Linux.
071
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 14/02/2026
I found loosely equivalent but much less interesting documents for Anthropic simonwillison.net/2026/Feb/13/anthr…
Anthropic's are much less interesting that OpenAI's. The earliest document from 2021 states:

    The specific public benefit that the Corporation will promote is to responsibly develop and maintain advanced Al for the cultural, social and technological improvement of humanity.

Every subsequent document up to 2024 uses an updated version which says:

    The specific public benefit that the Corporation will promote is to responsibly develop and maintain advanced AI for the long term benefit of humanity.
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 12/02/2026
New super-fast model from OpenAI today powered by their new Cerebras partnership - GPT-5.3-Codex-Spark It's 4-5x faster than GPT-5.3-Codex but the pelican isn't as good! Here's its pelican compared to full GPT-5.3 Codex, both on "medium" simonwillison.net/2026/Feb/12/codex…
Whimsical flat illustration of an orange duck merged with a bicycle, where the duck's body forms the seat and frame area while its head extends forward over the handlebars, set against a simple light blue sky and green grass background.Whimsical flat illustration of a white pelican riding a dark blue bicycle at speed, with motion lines behind it, its long orange beak streaming back in the wind, set against a light blue sky and green grass background.
211
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 12/02/2026
Since it did so well on the basic pelican I tried Gemini 3 Deep Think on my more detailed California Brown Pelican prompt as well - prompt in the alt text
Generate an SVG of a California brown pelican riding a bicycle. The bicycle must have spokes and a correctly shaped bicycle frame. The pelican must have its characteristic large pouch, and there should be a clear indication of feathers. The pelican must be clearly pedaling the bicycle. The image should show the full breeding plumage of the California brown pelican.

(It mostly nailed it but the pelican's legs are slightly detached from the body)
121
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 08/02/2026
In celebration of the 2026 breeding season ceramic artist Karen James made me a Kākāpō mug! simonwillison.net/2026/Feb/8/kakapo…
A simply spectacular sgraffito ceramic mug with a bold, charismatic Kākāpō parrot taking up most of the visible space. It has a yellow beard and green feathers.Another side of the mug, two cute grey Kākāpō chicks are visible and three red rimu fruit that look like berries, one on the floor and two hanging from wiry branches.
0102
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 05/02/2026
They had a most excellent editorial voice, as demonstrated by this "what's new" entry from December 10th 2020 about the height of Mount Everest simonw.github.io/cia-world-factbook…
December 10, 2020
Years of wrangling were brought to a close this week when officials from Nepal and China announced that they have agreed on the height of Mount Everest. The mountain sits on the border between Nepal and Tibet (in western China), and its height changed slightly following an earthquake in 2015. The new height of 8,848.86 meters is just under a meter higher than the old figure of 8,848 meters. The World Factbook rounds the new measurement to 8,849 meters and this new height has been entered throughout the Factbook database.
045
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 08/01/2026
I joined the Oxide and Friends annual predictions podcast episode this week - here are my 1, 3 and 6 year predictions for AI and LLMs (and Kākāpō parrots) simonwillison.net/2026/Jan/8/llm-pr…

    1 year: It will become undeniable that LLMs write good code
    1 year: We’re finally going to solve sandboxing
    1 year: A “Challenger disaster” for coding agent security
    1 year: Kākāpō parrots will have an outstanding breeding season
    3 years: the coding agents Jevons paradox for software engineering will resolve, one way or the other
    3 years: Someone will build a new browser using mainly AI-assisted coding and it won’t even be a surprise
    6 years: Typing code by hand will go the way of punch cards
325
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 01/01/2026
Made lemon pigs! 🍋 🐷
Two lemon pigs - a small one and a big one - lemons with toothpick feet and cloves for eyes and little ears and slit mouths with coins in them
031
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 31/12/2025
Here's my enormous round-up of everything we learned about LLMs in 2025 - the third in my annual series of reviews of the past twelve months simonwillison.net/2025/Dec/31/the-y… This year it's divided into 26 sections! This is the table of contents

    The year of “reasoning”
    The year of agents
    The year of coding agents and Claude Code
    The year of LLMs on the command-line
    The year of YOLO and the Normalization of Deviance
    The year of $200/month subscriptions
    The year of top-ranked Chinese open weight models
    The year of long tasks
    The year of prompt-driven image editing
    The year models won gold in academic competitions
    The year that Llama lost its way
    The year that OpenAI lost their lead
    The year of Gemini
    The year of pelicans riding bicycles
    The year I built 110 tools
    The year of the snitch!
    The year of vibe coding
    The (only?) year of MCP
    The year of alarmingly AI-enabled browsers
    The year of the lethal trifecta
    The year of programming on my phone
    The year of conformance suites
    The year local models got good, but cloud models got even better
    The year of slop
    The year that data centers got extremely unpopular
    My own words of the year
    That’s a wrap for 2025
03917
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 12/12/2025
Just pushed a new release of my LLM Python library and CLI tool for interacting with Large Language Models, adding support for GPT 5.1 and GPT 5.2 plus a few smaller bug fixes llm.datasette.io/en/stable/changelo…
0.28 (2025-12-12)

    New OpenAI models: gpt-5.1, gpt-5.1-chat-latest, gpt-5.2 and gpt-5.2-chat-latest. #1300, #1317

    LLM now requires Python 3.10 or higher. Python 3.14 is now covered by the tests.

    When fetching URLs as fragments using llm -f URL, the request now includes a custom user-agent header: llm/VERSION (https://llm.datasette.io/). #1309

    Fixed a bug where fragments were not correctly registered with their source when using llm chat. Thanks, Giuseppe Rota. #1316

    Fixed some file descriptor leak warnings. Thanks, Eric Bloch. #1313

    Fixed a deprecation warning for asyncio.iscoroutinefunction.

    Type annotations for the OpenAI Chat, AsyncChat and Completion execute() methods. Thanks, Arjan Mossel. #1315

    The project now uses uv and dependency groups for development. See the updated contributing documentation. #1318
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 10/12/2025
I made that because now if a Python package doesn't have good release notes I can click "copy diff" and paste into Claude/ChatGPT/Gemini and get genuinely useful release notes a few seconds later No need to even add a prompt, the models guess that's what you want from the diff
That tool showing the diff between Anthropic 0.74.0 and 0.74.1 - with a copy diff buttonPasting the diff into Claude Opus 4.5 returns:

This diff shows a patch release of the Anthropic Python SDK (0.74.0 → 0.74.1). The key change is:
Structured Outputs Beta Version Update
The beta header for structured outputs was updated from structured-outputs-2025-09-17t0 structured-outputs-2025-11-13 in the messages. py file. This affects the parse method in both the sync and async message clients.
Everything else is just standard version bumping across the metadata files (METADATA, RECORD, _version.py).
Is there something specific you'd like help with regarding this update?
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 10/12/2025
My most recent HTML tool is this one for loading code from the Python packaging index and displaying a diff between two versions - all running entirely in the browser, created using this prompt (which remixed a previous project) […] [Original post on fedi.simonwillison.net]
Screenshot of the PyPI Package Changelog tool, which lets you type in a name - here "LLM" - and then view the diff between different versions of that package Build a new tool pypi-changelog.html which uses the PyPI API to get the wheel URLs of all available versions of a package, then it displays them in a list where each pair has a "Show changes" clickable in between them - clicking on that fetches the full contents of the wheels and displays a nicely rendered diff representing the difference between the two, as close to a standard diff format as you can get with JS libraries from CDNs, and when that is displayed there is a "Copy" button which copies that diff to the clipboard
110
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 10/12/2025
Here's the table of contents, - by patterns I'm talking about things like hitting CORS-enabled APIs, using localStorage and URLs to store state, loading dependencies from CDNs and taking extensive advantage of rich copy and paste for both input and output to the tools you build

    The anatomy of an HTML tool
    Prototype with Artifacts or Canvas
    Switch to a coding agent for more complex projects
    Load dependencies from CDNs
    Host them somewhere else
    Take advantage of copy and paste
    Build debugging tools
    Persist state in the URL
    Use localStorage for secrets or larger state
    Collect CORS-enabled APIs
    LLMs can be called directly via CORS
    Don’t be afraid of opening files
    You can offer downloadable files too
    Pyodide can run Python code in the browser
    WebAssembly opens more possibilities
    Remix your previous tools
    Record the prompt and transcript
    Go forth and build
121
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 01/12/2025
A tiny TIL: if you are seeing "Error 153: Video player configuration error" on YouTube iframes embedded on your site a likely culprit is sending the "Referrer-Policy: same-origin" HTTP header (Django SecurityMiddleware sends this by default) Switching […] [Original post on fedi.simonwillison.net]
YouTube embed with a big ugly black square and text Watch video on YouTube, Error 153 Video player configuration error
231
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 29/11/2025
Out of curiosity I decided to try and run the numbers on how much Netflix you can watch for the energy cost of a ChatGPT prompt As far as I can tell it's between 5.1 and 10.2 seconds, depending on which end of the 2019 IEA Netflix energy usage […] [Original post on fedi.simonwillison.net]
In June 2025 Sam Altman claimed about ChatGPT that "the average query uses about 0.34 watt-hours".

In March 2020 George Kamiya of the International Energy Agency estimated that "streaming a Netflix video in 2019 typically consumed 0.12-0.24kWh of electricity per hour" - that's 240 watt-hours per hour at the higher end.

Assuming that higher end, a ChatGPT prompt by Sam Altman's estimate uses:

0.34 Wh / (240 Wh / 3600 seconds) = 5.1 seconds of Netflix

Or double that, 10.2 seconds, if you take the lower end of the Netflix estimate instead.

I'm always interested in anything that can help contextualize a number like "0.34 watt-hours" - I think this comparison to Netflix is a neat way of doing that.

This is evidently not the whole story with regards to AI energy usage - training costs, data center buildout costs and the ongoing fierce competition between the providers all add up to a very significant carbon footprint for the AI industry as a whole.
669
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 24/11/2025
Updated my post with this section about their improved protection against prompt injection attacks - definitely better, but the problem is that if an attacker gets 10 tries they'll still succeed 1/3rd of the time! […] [Original post on fedi.simonwillison.net]
Still susceptible to prompt injection #

From the safety section of Anthropic’s announcement post:

    With Opus 4.5, we’ve made substantial progress in robustness against prompt injection attacks, which smuggle in deceptive instructions to fool the model into harmful behavior. Opus 4.5 is harder to trick with prompt injection than any other frontier model in the industry:

    Bar chart titled "Susceptibility to prompt-injection style attacks"

On the one hand this looks great, it’s a clear improvement over previous models and the competition.

What does the chart actually tell us though? It tells us that single attempts at prompt injection still work 1/20 times, and if an attacker can try ten different attacks that success rate goes up to 1/3!

I still don’t think training models not to fall for prompt injection is the way forward here. We continue to need to design our applications
031
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 20/11/2025
Also notable is the new SynthID feature of the Gemini app - you can upload a photo to it and ask if it was generated by AI and it will detect the invisible SynthID watermarks added by Nano Banana Pro
Screenshot of a mobile chat interface displaying a conversation about AI image detection. The user has uploaded a photo showing two raccoons on a porch; one raccoon reaches inside a paper bag a bench while the other stands on the ground looking up at it. The conversation title reads "AI Image Creation Confirmed". The user asks, "Was this image created with ai?" The AI response, labeled "Analysis & 1 more", states: "Yes, it appears that all or part of this image was created with Google AI. SynthID detected a watermark in 25-50% of the image."
231
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 18/11/2025
For reference, here's a photo I took myself of a California brown pelican in breeding plumage a couple of weeks ago
A glorious California brown pelican perched on a rock by the water. It has a yellow tint to its head and a red spot near its throat.
220
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 18/11/2025
Here's the upgraded SVG pelican riding a bicycle prompt - I've added stricter requirements around the bicycle and specified that the pelican should be a California brown pelican in full breeding plumage
Generate an SVG of a California brown pelican riding a bicycle. The bicycle must have spokes and a correctly shaped bicycle frame. The pelican must have its characteristic large pouch, and there should be a clear indication of feathers. The pelican must be clearly pedaling the bicycle. The image should show the full breeding plumage of the California brown pelican.Gemini 3 Pro: It's clearly a pelican. It has all of the requested features. It looks a bit abstract though.GPT-5.1 The pelican is very round - it's a dumpy little fellow. Its body overlaps much of the bicycle. It has a lot of dorky charisma.Claude Sonnet 4.5. Oh dear. It has all of the requested components, but the bicycle is a bit wrong and the pelican is arranged in a very awkward shape.
404
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 15/11/2025
New release of my llm-anthropic plugin adding support for structured outputs (via the new official API - previously I faked it with tool calls) and Anthropic's web search feature github.com/simonw/llm-anthropic/rel…

    Support for Claude's new structured outputs feature for Sonnet 4.5 and Opus 4.1. #54
    Support for the web search tool using -o web_search 1 - thanks Nick Powell and Ian Langworth. #30
030
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 11/11/2025
I've now reached the "six coding agents in six terminal windows at once" phase of parallel agent delirium simonwillison.net/2025/Nov/11/six-c…
131
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 09/11/2025
For comparison, here are the pelicans riding bicycles drawn by GPT-5-Codex-Mini (the new model), GPT-5-Codex and full GPT-5 - all produced via the same hacked version of the Codex CLI tool
GPT-5-Codex-Mini. This is terrible. The pelican is an abstract collection of shapes, the bicycle is likewise very messed upGPT-5 Codex. It's a dumpy little pelican with a weird face, not particularly great but better than Mini.GPT-5: Much better bicycle, pelican is a bit line-drawing-ish but does have the necessary parts in the right places
020
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 09/11/2025
OpenAI partially released a new model yesterday called GPT-5-Codex-Mini No API access yet, but I did some truly horrible things to their Codex CLI app to get it to spit out this SVG of a pelican riding a bicycle
This is pretty bad. The bicycle is just about recognizable - a collection o f abstract lines and two circles - but the pelican is a weird little snow goblin tangled in a bundle of random lines hovering over the rest of the bike
220
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 06/11/2025
And here's an example of one of my code research prompts


    Create a performance benchmark and feature comparison report on PyPI cmarkgfm compared to other popular Python markdown libraries—check all of them out from github and read the source to get an idea for features, then design and run a benchmark including generating some charts, then create a report in a new python-markdown-comparison folder (do not create a _summary.md file or edit anywhere outside of that folder). Make sure the performance chart images are directly displayed in the README.md in the folder.
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 06/11/2025
Here's my research repo - each of the 13 folders is a different research project, and the README is automatically updated by an LLM to include summaries describing each one github.com/simonw/research?tab=read…
Screenshot of a README document with a right-side navigation panel; navigation menu shows: Filter headings search box, heading "Research projects carried out by AI tools" followed by project list: sqlite-query-linter (2025-11-04), h3-library-benchmark (2025-11-04), h3o-python (2025-11-03), wazero-python-claude (2025-11-02), datasette-plugin-skill (2025-10-24), blog-tags-scikit-learn (2025-10-24), cmarkgfm-in-pyodide (2025-10-22), python-markdown-comparison (2025-10-22), datasette-plugin-alpha-versions (2025-10-20), deepseck-ocr-nvidia-spark (2025-10-20), sqlite-permissions-poc (2025-10-20), minijinja-vs-jinja2 (2025-10-19), node-pyodide (2025-10-19)
100
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 04/11/2025
And in case you don't make it as far as the "miscellaneous tips" section, here's a bunch of lessons I learned about working with coding agents that I picked up along the way simonwillison.net/2025/Nov/4/datase…
When working on anything relating to plugins it’s vital to have at least a few real plugins that you upgrade in lock-step with the core changes. The tadd and radd shortcuts were invaluable for productively working on those plugins while I made changes to core.
Coding agents make experiments much cheaper. I threw away so much code on the way to the final implementation, which was psychologically easier because the cost to create that code in the first place was so low.
Tests, tests, tests. This project would have been impossible without that existing test suite. The additional tests we built along the way give me confidence that the new system is as robust as I need it to be.
Claude writes good commit messages now! I finally gave in and let it write these—previously I’ve been determined to write them myself. It’s a big time saver to be able to say “write a tasteful commit message for these changes”.
Claude is also great at breaking up changes into smaller commits. It can also productively rewrite history to make it easier to follow, especially useful if you’re still working in a branch.A really great way to review Claude’s changes is with the GitHub PR interface. You can attach comments to individual lines of code and then later prompt Claude like this: Use gh CLI to fetch comments on URL-to-PR and make the requested changes. This is a very quick way to apply little nitpick changes—rename this function, refactor this repeated code, add types here etc.
The code I write with LLMs is higher quality code. I usually find myself making constant trade-offs while coding: this function would be neater if I extracted this helper, it would be nice to have inline documentation here, this changing this would be good but would break a dozen tests... for each of those I have to determine if the additional time is worth the benefit. Claude can apply changes so much faster than me that these calculations have changed—almost any improvement is worth applying, no matter how trivial, because the time cost is so low.Internal tools are cheap now. The new debugging interfaces were mostly written by Claude and are significantly nicer to use and look at than the hacky versions I would have knocked out myself, if I had even taken the extra time to build them.
That trick with a Markdown file full of upgrade instructions works astonishingly well—it’s the same basic idea as Claude Skills. I maintain over 100 Datasette plugins now and I expect I’ll be automating all sorts of minor upgrades in the future using this technique.
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 01/11/2025
Just sent out the October edition of my sponsors-only monthly newsletter - you can pay me $10/month to send you less! Here's the table of contents simonwillison.net/2025/Nov/1/sponso…

    Coding agents and "vibe engineering"
    Claude Code for web
    NVIDIA DGX Spark
    Claude Skills
    OpenAI DevDay and GitHub Universe
    Python 3.14
    October in Chinese Al model releases
    Miscellaneous extras
    Tools I'm using at the moment
210
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 23/10/2025
Prompt -> Result tools.simonwillison.net/terminal-to…


    Build a new tool called terminal-to-html which lets the user copy RTF directly from their terminal and paste it into a paste area, it then produces the HTML version of that in a textarea with a copy button, below is a button that says "Save this to a Gist", and below that is a full preview. It will be very similar to the existing rtf-to-html.html tool but it doesn't show the raw RTF and it has that Save this to a Gist button

    That button should do the same trick that openai-audio-output.html does, with the same use of localStorage and the same flow to get users signed in with a token if they are not already

    So click the button, it asks the user to sign in if necessary, then it saves that HTML to a Gist in a file called index.html, gets back the Gist ID and shows the user the URL https://gistpreview.github.io/?6d778a8f9c4c2c005a189ff308c3bc47 - but with their gist ID in it

    They can see the URL, they can click it (do not use target="_blank") and there is also a "Copy URL" button to copy it to their clipboard

    Make the UI mobile friendly but also have it be courier green-text-on-black themed to reflect what it does

    If the user pastes and the pasted data is available as HTML but not as RTF skip the RTF step and process the HTML directly

    If the user pastes and it's only available as plain text then generate HTML that is just an open <pre> tag and their text and a closing </pre> tag
Terminal to HTML app. Green glowing text on black. Instructions: Paste terminal output below. Supports RTF, HTML or plain text. There's an HTML Code area with a Copy HTML button, Save this to a Gist and a bunch of HTML. Below is the result of save to a gist showing a URL and a Copy URL button. Below that a preview with the Claude Code heading in ASCII art.
100
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 22/10/2025
Asynchronous coding agents are the fastest and safest route to running coding agents in a sandbox without constant supervision
The best sandboxes run on someone else's computer

Claude Code for Web
OpenAl Codex Cloud
Gemini Jules
ChatGPT & Claude code Interpreter
110
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 22/10/2025
Just for fun, I had Claude Code figure out how to run the ~2001-era Perl and C SLOCCount program in WebAssembly in the browser, complete with a UI for counting source code lines from pasted text, a GitHub repository or a zip file […] [Original post on fedi.simonwillison.net]
221
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 21/10/2025
It's neat to see them encourage developers to add ARIA tags to pages though, an "agent" can be thought of as effectively another form of assistive technology
There was one other detail in the announcement post that caught my eye:

    Website owners can also add ARIA tags to improve how ChatGPT agent works for their websites in Atlas.

Which links to this:

    ChatGPT Atlas uses ARIA tags---the same labels and roles that support screen readers---to interpret page structure and interactive elements. To improve compatibility, follow WAI-ARIA best practices by adding descriptive roles, labels, and states to interactive elements like buttons, menus, and forms. This helps ChatGPT recognize what each element does and interact with your site more accurately.

A neat reminder that AI "agents" share many of the characteristics of assistive technologies, and benefit from the same affordances.
323
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 18/10/2025
Here's my vibe-coded tool for displaying the Responses JSON returned from a deep research API call in a more readable way: tools.simonwillison.net/deep-resear… - built by Claude Code in this session […] [Original post on fedi.simonwillison.net]
Dashboard screenshot showing metrics at top: 17 Thinking Steps, 45 Searches, 24 Pages Visited, 12 Code Executions, 180 Total Steps. Below is a blue "Thinking" section with brain emoji containing text "**Researching orchestrions**" followed by a paragraph: "I'm considering a deep dive into specific orchestrions, particularly targeting places like museums. The idea is to gather data on surviving orchestrions and produce a structured list in a JSON format. Each entry will likely include details like city, country, venue, and notes about their history and significance. I realize this could be a challenging task, as orchestrions are quite rare. The goal is to compile a comprehensive overview, so I need to identify reliable sources of information." At bottom is a beige search box with magnifying glass icon showing: Search: "surviving orchestrion" locations
000