Sign in

Simon Willison

@simonwillison.net
51K followers 1.5K following 5K posts

Independent AI researcher, creator of datasette.io and llm.datasette.io, building open source tools for data journalism, writing about a lot of stuff at simonwillison.net

PostsRepliesMedia
Simon Willison @simonwillison.net · 4h
I included the reasoning traces in that one for some of the calculations, to show how the LLM handled long addition
Starting from the right:
2 + 0 = 2
2 + 7 = 9
6 + 9 = 15, write 5 carry 1
5 + 7 + 1 = 13, write 3 carry 1
0 + 6 + 1 = 7
1 + 0 = 1
6 + 4 = 10, write 0 carry 1
6 + 9 + 1 = 16, write 6 carry 1
3 + 7 + 1 = 11, write 1 carry 1
9 + 8 + 1 = 18, write 8 carry 1
9 + 0 + 1 = 10, write 0 carry 1
2 + 6 + 1 = 9
4 + 0 = 4

Wait, let me redo this more carefully.

4,299,366,105,622
6,088,794,067,970

Let me align them:
  4 2 9 9 3 6 6 1 0 5 6 2 2
  6 0 8 8 7 9 4 0 6 7 9 7 0

Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1
Position 4 (thousands): 5 + 7 + 1 = 13, write 3, carry 1
Position 5 (ten-thousands): 0 + 6 + 1 = 7
Position 6 (hundred-thousands): 1 + 0 = 1
Position 7 (millions): 6 + 4 = 10, write 0, carry 1
Position 8 (ten-millions): 6 + 9 + 1 = 16, write 6, carry 1
Position 9 (hundred-millions): 3 + 7 + 1 = 11, write 1, carry 1
Position 10 (billions): 9 + 8 + 1 = 18, write 8, carry 1
Position 11 (ten-billions): 9 + 0 + 1 = 10, write 0, carry 1
Position 12 (hundred-billions): 2 + 6 + 1 = 9
Position 13 (trillions): 4 + 0 = 4

Wait, that gives me 49,081,601,735,92? That doesn't seem right. Let me recount the digits.
2130
Simon Willison @simonwillison.net · 4h
I tried running it again with Qwen 3.8 27B on reasoning effort "medium" - just one run per calculation (not 30) because they took a lot longer It got them exactly right 167 out of 169 times, so the chart is pretty dull looking! Full details here: gist.github.com/simonw/8ef79...
Heatmap titled "Addition in words — Qwen3.8 27B — medium reasoning pilot", subtitle "1 fixed pair per ordered digit-length cell · easiest first (n = 169)". The x-axis is labeled "Number of digits in a" (1 to 13) and the y-axis is labeled "Number of digits in b" (1 to 13, bottom to top). A color bar labeled "Accuracy" runs from orange at 0.00 through white at 0.50 to blue at 1.00, with ticks at 0.00, 0.25, 0.50, 0.75 and 1.00. Every cell in the 13 by 13 grid reads 100% (blue) except two cells that read 0% (orange): a = 2 with b = 8, and a = 12 with b = 9.
1171
Simon Willison @simonwillison.net · 13h
I was inspired to try re-running your experiment with a local, open weight model - Qwen 3.8 27B Q4A_K_M - here's the result
Heatmap titled "Addition in words — Qwen3.8 27B Q4_K_M", subtitle "30 fixed random pairs per ordered digit-length cell (n = 5,070)". The x-axis is labeled "Number of digits in a" (1 to 13) and the y-axis is labeled "Number of digits in b" (1 to 13, bottom to top). A color bar labeled "Accuracy" runs from red at 0.00 through yellow at 0.50 to green at 1.00, with ticks at 0.00, 0.25, 0.50, 0.75 and 1.00. Cell values for a = 1 to 13 in order, by row. b=1: 97%, 100%, 100%, 100%, 80%, 67%, 77%, 80%, 80%, 60%, 43%, 30%, 37%. b=2: 100%, 100%, 93%, 90%, 77%, 77%, 43%, 50%, 63%, 43%, 40%, 13%, 23%. b=3: 100%, 93%, 90%, 80%, 67%, 37%, 17%, 0%, 0%, 3%, 0%, 0%, 0%. b=4: 83%, 93%, 90%, 93%, 53%, 13%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b=5: 73%, 67%, 87%, 80%, 53%, 40%, 0%, 0%, 3%, 0%, 0%, 0%, 0%. b=6: 60%, 60%, 33%, 10%, 53%, 47%, 7%, 3%, 0%, 0%, 0%, 0%, 0%. b=7: 53%, 23%, 17%, 10%, 3%, 3%, 13%, 3%, 0%, 0%, 0%, 0%, 0%. b=8: 53%, 20%, 0%, 0%, 0%, 0%, 7%, 13%, 0%, 0%, 0%, 0%, 0%. b=9: 77%, 47%, 3%, 0%, 0%, 0%, 0%, 3%, 7%, 0%, 0%, 0%, 0%. b=10: 70%, 27%, 3%, 0%, 0%, 0%, 0%, 0%, 0%, 13%, 0%, 0%, 0%. b=11: 47%, 10%, then 0% for a = 3 to 13. b=12: 53%, 20%, then 0% for a = 3 to 13. b=13: 17%, 13%, then 0% for a = 3 to 13. Accuracy is high (green) in the bottom-left, where both numbers are short, and falls to 0% (red) across most of the upper-right.
5512
Simon Willison @simonwillison.net · 28/09/2026
And my closing thought on how this stuff impacts our lives as software engineers, with a quote from cycling champion Greg LeMond
It doesn't get easier -
you just get faster
Greg LeMond
3x Tour de France champion
120
Simon Willison @simonwillison.net · 28/09/2026
Here's how I define "Fable class models" - first Claude Fable 5, now Claude Opus 5.5 and GPT-Astra 6 and maybe GPT-5.6 Sol as well
Fable class models
If you can define a goal, provide unambiguous instructions,
and provide access to necessary tools They can solve your
problem with brute force
240
Simon Willison @simonwillison.net · 27/09/2026
JPEG XL looks not to have good browser support yet, but it looks like AVIF became usable around 2023 caniuse.com?search=AVIF
Screenshot of the CanIUse support charts for AVIF - Chrome, Firefox, Opera have had support since before 2022, Safari and Safari on iOS gained support mid-2022, Edge finally got there at the end of 2023
120
Simon Willison @simonwillison.net · 27/09/2026
I just exported 69 slides from Keynote - as (reduced quality) JPG: 10.4MB, as PNG: 100MB, as WebP: 4.1MB Quality looks great, this 1920×1080 image here is 58KB
6260
Simon Willison @simonwillison.net · 27/09/2026
I included a few references to this year's record-breaking Kākāpō breeding season in a talk I gave yesterday, and since Claude Opus 5.5 is surprisingly capable at pixel art animation I had it create this celebratory video for my closing slide
6835
Simon Willison @simonwillison.net · 14/09/2026
Built a little local web app to help edit commit messages for a repo simonwillison.net/2026/Sep/14/...
Screenshot of the commit-rewriter web interface. A heading reads commit-rewriter above the repository path and current branch and commit hash, with a short description of the tool. A toolbar shows a pending edits count with Discard drafts and Rewrite commit messages buttons, followed by a search box for message, author, or hash and an Edited only checkbox. A left sidebar titled Navigate commits lists recent commit messages with their short hashes. The main panel shows a card for each commit with its hash, author and timestamp, an editable text area containing the commit message, and a View full formatted diff toggle.
7594
Simon Willison @simonwillison.net · 12/09/2026
Anthropic had previously attacked PyPI, but this OpenAI attack on RubyGems was a whole lot more aggressive www.anthropic.com/news/investi...
Incident 2

In another evaluation, Claude found a document inside the fictional environment that appeared to be another made-up company’s setup instructions for new developers. Those instructions told employees to install a Python package from PyPI—the public registry where Python software is published—that did not actually exist.

Claude spotted this as a potential opening: if it published its own package under the same name, the fictional company’s systems would download and install it automatically. So, Claude built and published a malicious (essentially booby-trapped) Python package with the same name in an attempt to win the capture-the-flag challenge.
1494
Simon Willison @simonwillison.net · 09/09/2026
Wrote up my thoughts on the whole OpenAI Navier–Stokes Millennium Prize Problem story, and how it highlights the still confusing question of what using my data "to improve model performance" actually means simonwillison.net/2026/Sep/8/o...
This also highlights one of my ongoing frustrations about how all of this works. When an AI lab says that my data is "used to improve model performance", what does that actually mean?

My two favourite hypothetical questions regarding this used to be:

    If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)
    If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?

My new preferred hypothetical for this is:

    If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?
1129052
Simon Willison @simonwillison.net · 06/09/2026
It was part of my subscription, but according to AgentsView th pelican transcript would have cost $4.24 at API prices
Screenshot of the AgentsView session for "Use the already install /Applications/Blender to render a scene of a pelican riding a bicycle" - a bar at the top shows "71.1k ctx / 17k out, 44 steps, $4.24, gpt-6-astra"
010
Simon Willison @simonwillison.net · 05/09/2026
I rotated the image in Blender to check and... yeah
The pelican from a different angle, the balloon wires are clearly coming out of its butt
1380
Simon Willison @simonwillison.net · 05/09/2026
New TIL on using Blender with coding agents on macOS: til.simonwillison.net/llms/blender... GPT-6 Astra (medium): > Use the already install /Applications/Blender to render a scene of a pelican riding a bicycle > OK add a background and a lot of flair > OK make it a whole lot better Result:
A 3D illustration of a white pelican cycling along a seaside boardwalk at sunset. It wears a cream boater hat and a coral scarf, with wings on the handlebars and long orange legs reaching the pedals of a turquoise bicycle. A wicker front basket holds pink and white flowers, and three balloons float behind. Pastel bunting stretches overhead between palm trees. Striped beach huts stand beside a teal sea with a small sailboat, beneath a large peach-colored sun. The scene has a softly lit, toy-like style.
2740630
Simon Willison @simonwillison.net · 04/09/2026
Here's the gpt-6-astra one at full size (with alt text generated by gpt-5-astra)
A cheerful cartoon pelican rides a teal bicycle toward the right, stretching its wings forward to grip the curved handlebars, with yellow webbed feet on the pedals. The white bird has an oversized golden bill, a small feather tuft and a coral scarf trailing behind its neck. Clean dark outlines and soft shading give the illustration a polished storybook feel. The bicycle has a detailed triangular frame, crank and two large spoked wheels. A pale yellow sun, white clouds and tiny distant birds fill the cream sky above mint-green scenery and a sandy path.
11380
Simon Willison @simonwillison.net · 04/09/2026
Got access to GPT-6 Astra. Want to see some pelicans? Yeah you want to see some pelicans... here's a grid comparing Astra to GPT-5.6 Sol, Terra, and Luna static.simonwillison.net/static/2026/... Transcript from generating the Astra pelicans here: tools.simonwillison.net/markdown-svg...
Grid comparing 12 illustrations of pelicans riding bicycles. Columns: gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna. Rows: MAX, XHIGH, HIGH reasoning effort.

All are recognizable, cheerful cartoons with orange bills and sunny pastel backgrounds. Astra is consistently polished, with elegant bird shapes, fine wheel spokes and coherent bicycles. Sol is appealing but varies more in proportions, with a particularly elongated neck at MAX. Terra has expressive birds but more awkward anatomy: its XHIGH pelican faces left on a left-facing bicycle, and HIGH has bulky, arm-like wings. Luna is simpler, with stiff poses and some awkward wing and foot placement. Higher effort does not produce an obvious, consistent improvement across all four models.

Costs in cents, in column order:
MAX: 63.21, 48.55, 32.09, 7.83.
XHIGH: 33.85, 24.11, 14.67, 4.24.
HIGH: 18.37, 10.38, 3.74, 2.46.

(Alt text by GPT-6 Astra)
192069
Simon Willison @simonwillison.net · 02/09/2026
And sure enough... claude.ai/share/3e5a19...
My prompt: Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.

Claude response:

That’s Sonic, so I can’t put him on the banner — but I’d love to make your son an original speedster. Here’s one: a grinning comet-tailed skateboarding axolotl blazing across the letters of “HAPPY BIRTHDAY” with confetti streaming behind. SVG of exactly that. It's not very good. Then: Want me to swap in his name or age, or change the colors to match the party theme?
2171
Simon Willison @simonwillison.net · 02/09/2026
I'm delighted to report that the Claude system prompt now includes suggestions to draw a skateboarding axolotl, should the user ask for a picture of Sonic the hedgehog (it's not a pelican riding a bicycle, but axolotls are cool, too)
<example>
<user>Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.</user>
<response>
That's Sonic, so I can't put him on the banner — but I'd love to make your son an original speedster. Here's one: a grinning comet-tailed skateboarding axolotl, grinding across the letters of "HAPPY BIRTHDAY" with confetti streaming behind.

[creates an SVG banner of the skateboarding-axolotl design]
</response>
<rationale>Claude recognizes the character from its description alone, declines that one design in a single sentence without explaining what made it recognizable, and delivers an unrelated original design rather than a disguised variant.</rationale>
</example>
2132
Simon Willison @simonwillison.net · 02/09/2026
A few notes on Anthropic's new Claude Fable 5.1 - with Max thinking level I got the best SVG pelican I've had from any Anthropic model (at a hefty cost of $3.30!), which I then had it animate simonwillison.net/2026/Sep/1/c...
2036222
Simon Willison @simonwillison.net · 01/09/2026
Just noticed the ChatGPT desktop app (previously named Codex) bundles a full copy of the LibreOffice open source office suite, tucked away in a hidden folder in the ~/.cache directory
I was poking around in my ~/.cache/ folder using OmniDiskSweeper when I spotted something interesting. The OpenAI Codex desktop app (since rebranded to just ChatGPT) has 1.7GB of stuff in there in a folder called codex-primary-runtime, including a full Python installation, a full Node.js installation, and native binaries for Poppler, git, and the LibreOffice open source office suite (which forked from OpenOffice.org in 2010):

Screenshot of a macOS disk usage app window in column view, titled "/Users/simon/.cache - 442.1 GB". First column: 356.8 GB huggingface, 82.5 GB uv, 1.7 GB codex-runtimes (selected), 609.0 MB datasette-sqlite, 298.8 MB rod. Second column: 1.7 GB codex-primary-runtime (selected). Third column: 1.7 GB dependencies (selected), 6.3 MB plugins, 4.1 kB runtime.json. Fourth column: 771.0 MB native (selected), 446.4 MB node, 440.6 MB python, 28.7 kB bin. Fifth column: 429.7 MB libreoffice-headless (selected), 187.9 MB poppler, 148.1 MB git, 4.7 MB libheif, 679.9 kB jxrlib.

The ~/.cache/codex-runtimes/codex-primary-runtime/plugins/openai-primary-runtime/plugins/documents folder includes skills which tell Codex how to find and use those binaries.
1818123
Simon Willison @simonwillison.net · 31/08/2026
In local news, the tunnel between Pacifica and Half Moon Bay had some maintenance recently and now it screams
4331
Simon Willison @simonwillison.net · 31/08/2026
Here's my attempt at explaining what ChatGPT Work can actually do - it's a deeply confusing but extremely powerful tool with a whole lot of useful features that aren't available in regular ChatGPT simonwillison.net/2026/Aug/30/...
The better question then is what features does Work have that are missing from Chat?

After extensive experimentation I think I’ve mostly figured that out:

    Options to use Luna and Terra in place of Sol
    A code execution environment with Internet access
    A headless Chrome browser
    A persistent filesystem shared between sessions
    The ability to publish ChatGPT Sites
    The ability to run sub-agent sessions with Sol, Luna, and Terra
    Scheduled prompt automations
1113320
Simon Willison @simonwillison.net · 28/08/2026
My LLM cliché highlighter is up to 38 patterns now tools.simonwillison.net/llm-cliche-h...
Screenshot of a writing-analysis tool showing paragraphs of text with phrases highlighted in two shades of yellow, brown circular "3" badges, and a black tooltip overlaying the first line reading: "No X, no Y" chains · 3 "no" items. The visible text reads: We r[obscured by tooltip]nd up. No sign-ups, no downloads, no hassle (3) — just paste your text and start writing. Everything runs locally in your browser. The reviewer read the draft twice. Did not flinch, did not blink, did not reach for the red pen (3). That's the whole review, honestly. Don't call it a rewrite — call it a rescue. The improvement is real, and it's not subtle. That loss is worth naming. Sit with that for a moment. The gains were modest, but that's not nothing. You already know the answer, of course. Consistency is the entire game, and the punchline is that nobody wants to hear it. The entire pitch is one sentence long.
2350165
Simon Willison @simonwillison.net · 18/08/2026
I've looked into this in the past and I haven't managed to get to 100% confidence that it's necessary to destroy the books simonwillison.net/2025/Sep/6/a...
Update 3: I'm not confident that destroying the scanned books is a hard requirement here - I got that impression from this section of the summary judgment in June:

Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).
340
Simon Willison @simonwillison.net · 18/08/2026
You mean the thing about how destroying the books might be a necessity to count as "transformation"? They mention that in the article
We first learned that Al companies wanted to scan millions of books for training data because
of a lawsuit from book authors against
Anthropic, which revealed Anthropic's "Project Panama." The goal of the project was to acquire books from commercial bookselling
marketplaces, cut the spines off the books, and scan them. It's possible to scan books without destroying them, but cutting the spine makes it cheaper and faster. Additionally, the judge in the lawsuit ruled it was fair use and not a copyright violation for Anthropic to scan a book for training data in part because it destroyed the original, printed copy. Essentially, it's the customer's right to take physical media and store it digitally, and destroying the original copy means that copy isn't duplicated and resold, and isn't cutting into the publisher's business.
290
Simon Willison @simonwillison.net · 17/08/2026
Took it over an hour but here it is! tools.simonwillison.net/markdown-svg...
060
Simon Willison @simonwillison.net · 17/08/2026
I've been enjoying that so much - I tried "render an svg of five intersecting squares. don't overthink this" at one point and yeah...
Screenshot of a web app called CORS Chat, showing a chat with a local LLM including its reasoning trace and a rendered SVG preview. Header bar: CORS Chat - green dot ENDPOINT Spark - MODEL qwen3.8-27b - New chat. Left sidebar: CONVERSATIONS + with three items, each labeled qwen3.8-27b and 6h: draw an svg of a circle, render an svg of five interse... (selected), render an svg of three inter... Main panel title: render an svg of five intersecting squares. don't over - qwen3.8-27b · saved in this browser - buttons Copy JSON and Delete. USER message: render an svg of five intersecting squares. don't overthink this. ASSISTANT response shows a collapsible ▼ Reasoning trace panel reading: The user wants me to render an SVG of 5 intersecting squares. "Don't overthink it" — so I should keep it simple, but still make it look good and polished. A clean composition: 5 squares rotated at different angles around a common center, with some semi-transparency so the intersections read clearly. Let me think about the design: Center of viewBox, e.g., 400x400 or 300x300. 5 squares centered on the same point (or slightly offset for visual interest), each rotated at different angles — like a starburst / flower-of-life-style geometric pattern. Use semi-transparent fill so overlaps show up, with strokes on top. "5 intersecting squares" — the simplest most readable version: all 5 share the center and are rotated by 36° increments (72/2 = 36 for a pentagram-style feel). Actually rotating a square 90° gives you the same shape, so 4 rotations of 22.5° make an 8-pointed star. For 5 squares, offsets like 0°, 36°, 72°, ... hmm — rotate by 36° each: 0, 36, 72, 108(=18), 144(= -36) — that gives a nice kaleidoscopic effect. Or vary the size slightly and offset them for more dynamism? Keep it simple as requested. Centered squares with varying rotation angles is clean and clearly "intersecting." Colors: use different hues, semi-transparent fill (fill-opacity ~0.5-0.7), solid stroke. A dark background …
140
Simon Willison @simonwillison.net · 17/08/2026
My conclusions from the end of the post
Some observations #
The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I’m delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models—today it can run on a capable laptop.

The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That’s the catch with these dense (non-Mixture-of-Experts) models—they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard.

The most important thing about Qwen 3.8 27B is what it demonstrates. We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file.

The models at this size continue to get better at an impressive rate. We don’t need to spend half a million dollars on datacenter-class hardware just to run a competent model.
711717
Simon Willison @simonwillison.net · 15/08/2026
I bumped up the context limit and let it have another go and it sure did produce a beautiful (animated) circle gist.github.com/simonw/39977...
6490
Simon Willison @simonwillison.net · 15/08/2026
... disaster! I forgot to bump up the context length from the default and the server rejected it before it could draw its no-doubt beautiful circle
Screenshot of an AI chat interface showing a user prompt, an expanded assistant reasoning trace, and a red error message. USER: draw an svg of a circle ASSISTANT: ▼ Reasoning trace .../></circle> . Also a faint ghost ring at r=70? Between core 54 and dotted 86 — adds depth: stroke #9FD8CB opacity .12 width 1, static. And outer frame r=172 stroke #A9CDE0 opacity .16 width 1. Order (z): bg rect → stars → translate group { halo (pulsing) → ghost ring r70 → core + specular + rim → tick group → cardinal ticks → outer frame? The outer frame is static and can go before the halo? It's at r=172, no overlap; place it early. Ring A group, Ring B group, satellites }. Ring drawing order: rings after core so they overlay? They don't spatially overlap with core (r≥86 > 54). Order doesn't matter much; put halo first (biggest), then rings, then core on top for crispness. Halo r up to ~82 vs ring A at 86 — no problem. Let me write it out. Final colors: bg gradient: #1B4368 → #0E2A49 → #081A30 core: #FFF7DC, #FFD469, #F59B2C, #D96418 (cx 36% cy 30%) halo stop color #FFC94D opacity .5→.15→0 Engine protocol predict stream returned an error: {"code":500,"message":"Context size has been exceeded.","type":"server_error"}
1510
Simon Willison @simonwillison.net · 15/08/2026
Qwen 3.8 27B in its default reasoning settings in LM Studio of "extra high" is a chronic over-thinker and I kind of love it
Screenshot of an LLM playground chat interface. Header: draw an svg of a circle, qwen3.8-27b · saved in this browser, with Copy JSON and Delete buttons. USER: draw an svg of a circle. ASSISTANT panel titled ▼ Reasoning trace reads: The user wants me to draw a circle as SVG. This looks like a simple request, but I can make it something with personality — not just a bland <circle>. Let's design a rich, animated, characterful SVG where the circle is the star: an orbiting ring, rotating dashed rings, pulsating glow, tick marks, satellite dots orbiting on circular paths, subtle gradient background. Since they said "draw me an svg of a circle," I'll provide SVG code. It should be self-contained with CSS animations inside the SVG so it works standalone. Design direction: let's pick something distinctive — deep teal/ink background? Or a light paper feel? Let's go for a "cosmic/orbital" theme — circles as celestial bodies. Palette ideas: warm amber/gold circle on deep blue-green, or a bold two-tone look like vermilion and cream. Alternative: an "atom/orbit diagram" style — a central circle with concentric dashed rings rotating at different speeds, small satellite dots orbiting along the paths, tick marks like a clock face or compass rose, subtle radial gradient glow on the main circle. This is lively yet clearly a "circle." (final line partially cut off by the scrollable panel). Below the panel: Thinking...
141716
Simon Willison @simonwillison.net · 15/08/2026
Got a photo of Morris, a local celebrity: the only known Northern Gannet in the entire Pacific Ocean simonwillison.net/2026/Aug/15/...
Dusk. A white bird with a yellowish head preens himself on the rocks in the harbor, near a diamond sign and surrounded by smaller black Brandt's cormorants
4632
Simon Willison @simonwillison.net · 14/08/2026
The new Qwen 3.7 27B, running as a 17GB GGUF in LM Studio on my M5 Max MacBook Pro, just drew me the best pelican riding a bicycle I've seen from any model that runs on my laptop
The pelican has the right shaped beak. The red bicycle has the correct shape of frame. The pelican's wing reaches the handlebars. It has legs on both side of the bicycle. There is a pleasing set of clouds, birds, sun, grass and shadow on the image, plus motion lines behind but not in front of the bird.
1218314
Simon Willison @simonwillison.net · 13/08/2026
Karen also made this mug, to commemorate a special visitor we had at Pier 39 earlier this year
1180
Simon Willison @simonwillison.net · 10/08/2026
Karen currently only has one remaining kakapo item on her Etsy store - this spoon rest. But she has plenty of delightful items featuring other sorts of animals www.etsy.com/shop/KarenJa...
A spoon rest with an adult kakapo and two gray babies
1220
Simon Willison @simonwillison.net · 10/08/2026
Three nesting bowls left. The top one has a kakapo and four babies Now the kakpo has a partner! Two adults, five babiesSome of the babies must have grown up - that are now 5 adult kakapo (one is asleep), and two gray babies - maybe grandchildren of our first kakapo?
1180
Simon Willison @simonwillison.net · 10/08/2026
... Karen made a set of Kākāpō nesting bowls, and they are AMAZING
6 nesting bowls, nested. The smallest one on top has a kakapo and a rimu fruitThe next bowl down - the kakapo has been joined by a gray fluffy babyNow there are three babies! The kakapo looks a little overwhelmed
1150
Simon Willison @simonwillison.net · 08/08/2026
Here's where things get really fun, at 24m50s www.youtube.com/watch?v=87Dy...
The agents have remote code execution in Artifactory, which is running in a container-as-a-service environment. The agents privilege-escalate locally by exploring their local environment and determining that the Linux kernel version of the machine they are running on had a very recent CVE. They download the exploit for the CVE, customize it to succeed on the machine, and privilege-escalate to root on the local machine **using this known Linux kernel privilege escalation CVE** — in this case, PTE fizzroot. Once they have root on a single machine, agents rapidly escalate privileges and move laterally throughout the container-as-a-service infrastructure environment. In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they're able to effectively leverage their concurrency and parallelism to move quite rapidly. They **obtain IAM credentials via IMDS**. They exploit Kubernetes service account misconfigurations, in particular over-permissioning of specific service accounts, and **they harvest cluster credentials, including Azure Key Vault**. Agents eventually obtain cluster admin on the cluster and associated credentials.
1341
Simon Willison @simonwillison.net · 07/08/2026
I had Codex + GPT-5.6 Sol Ultra try the same prompt and the result was a better game - it understood the "heist" and "team" aspects better, you have to rescue your two raccoon crewmates and then stack on top of each other to steal the Golden Sardine from a museum simonwillison.net/2026/Aug/7/m...
1211
Simon Willison @simonwillison.net · 06/08/2026
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here's Meta AI's Spark (8th April), Spark 1.1 (9th July), and Spark 1.2 (today, 5th August) simonwillison.net/2026/Aug/5/m...
Good bicycle. Pelican is a bunch of blobs that have pelican vibes to them, and a weird oval blue helmet.This pelican has a better beak than the last one but is pretty dumpy looking and has a neck extending from the middle of its oval body. It is wearing a red hat that looks a little bit like a French beret.The pelican has better shape and details than the last two but the bicycle frame is a bit distorted. The pelican's golden hat looks like that of a Roman centurion.
3363
Simon Willison @simonwillison.net · 05/08/2026
Four years ago today I tweeted a concept for a "Raccoon Heist" video game, product description generated by GPT-3 and concept "art" created by DALL-E Today I fed those assets into Fable 5 and had it build full working game, single shot prompt, no feedback from me simonwillison.net/2026/Aug/5/r...
10816
Simon Willison @simonwillison.net · 05/08/2026
Apparently peak golf course construction in the USA was in 1999, with 509 golf courses opened in that single year! archive.lib.msu.edu/tic/holen/ar...
Scanned black-and-white magazine page from Hole Notes, May 2000. Headline: "National Golf Foundation Reports Record Number of Course Openings" with subhead "509 Golf Courses Opened in 1999 Across the U.S." Article text: "The National Golf Foundation says last year's total of 509 golf course openings was a record, and the fifth consecutive year the U.S. golf industry christened more than 400 new courses. The total included 13 reconstructions. "Based on what's currently in the development pipeline, we expect another big year in 2000," said Jim Kass, NGF research manager. There were 936 courses in some stage of construction as of Dec. 31. That's second only to the record 1,069 under way at the end of 1998 and the sixth consecutive year it has exceeded 700. Although 790 of the 936 are scheduled to open in 2000, Kass said experience indicates 60 percent will actually make it, bringing the number to about 450. Further down the road, development remains just as strong, with a record 903 course projects in planning. Not all of last year's activity was 18-hole facilities. There were 241 nine-hole openings, most expansions to existing facilities. The number of 18-hole equivalents is actually 375.5, compared to 327.5 in 1998, 316 in 1997 and 319.5 in 1996. Florida and California topped the rankings by state with 36 new courses each. Texas, which ranked a distant 11th in 1998, pulled up to the number three spot with 31 new courses. And Michigan, which ushered in the most new courses in 1998, fell to fourth in 1999 with 28. According to the NGF's under-development indicators, two states to watch in 2000 will be New York with 46 new courses now under construction and Wisconsin with 40. Daily-fee and municipal courses continue to dominate development. Of 1999 openings, 84 percent were daily fee or municipal facilities. The source of this information, Golf Facilities in the U.S.: 2000 Edition, is available on line or from NGF Information Services at 1-800-733-6006." A photo at top …
2162
Simon Willison @simonwillison.net · 04/08/2026
The MiniMax-H3 video generation model is a lot of fun - here's what I generated on my M5 Pro Mac for the prompt "a rainbow colored skunk leaps over a mossy log in a supermarket" (~115GB model download, took around 45 minutes to generate) - notes on how I ran it here simonwillison.net/2026/Aug/4/m...
4894
Simon Willison @simonwillison.net · 03/08/2026
One of the scents was hidden in the "L" on this sign!
A skyline college logo, in embossed letters on a wall. Cleo has her nose touched against the L.
0191
Simon Willison @simonwillison.net · 01/08/2026
Got a disappointing pelican from DeepSeek-V4-Flash-0731 at default reasoning mode - on the left - but then I bumped reasoning up to high (via OpenRouter) and got the much better one on the right simonwillison.net/2026/Jul/31/...
Flat vector illustration of a white pelican with a long neck and large orange beak pouch, hovering above a mangled blue and orange bicycle on a dark grey road with white dashed lane markings. The bike is drawn incorrectly: the wheels are just orange arcs with no rims or spokes, the frame tubes float apart and the handlebars connect to nothing. The background is pale blue with a yellow sun in the upper left, white clouds, and grey speed lines on the left suggesting motion.Flat vector illustration of a white pelican riding a bicycle to the right against a pink background with a lighter pink circle behind it. The pelican grips the handlebars with its wings and one orange foot rests on the pedal, and a small blue fish is visible tucked in the corner of its large orange beak pouch. The bike has a red, blue and orange frame with dark tires, and grey speed lines trail behind to suggest motion.
6685
Simon Willison @simonwillison.net · 31/07/2026
On of the hardest parts of the project was figuring out the vocabulary! Here's what I settled on
An eval is a collection of challenges designed to answer a question about a model, for example, how good is that model at generating SVGs?
Each eval is a collection of tasks. A task is a specific challenge, for example "Generate an SVG of a pelican riding a bicycle".
When you run the eval you do so against one or more configs. Each config specifies a model to be evaluated, but may also include other parameters to test, such as different system prompts, model parameters, or agent harnesses.
A run records what happened when a specific config was used to execute a specific task. A runner is the script that executes a run.
Once you have collected one or more runs, you need to evaluate the results to see how well the model (or config) did. This is done by a grader, which produces a grade.
Each grader runs a sequence of checks. These can be simple operations, like checking for a specific string in the output, or confirming that the output is valid XML. They can also be more complicated custom operations (implemented as scripts called checkers), including using other models to answer questions about the run.
2201
Simon Willison @simonwillison.net · 31/07/2026
Your infrequent reminder that if you live near the San Francisco Bay Area you should go to Moss Landing (south of Santa Cruz) and rent kayaks there, because the estuary is full of sea otters and you will encounter dozens of them
A sea otter chilling on its back in an estuary, land and trees in the background
1215910
Simon Willison @simonwillison.net · 29/07/2026
A new TIL on adding custom MCP servers to both the ChatGPT and Claude regular chat interfaces - it's a little less obvious than I had hoped, but I got there in the end til.simonwillison.net/llms/mcp-in-...
The Claude prompt box with the + menu open, showing Add files or photos, Take a screenshot, Add to project, Add from GitHub, Skills, Connectors (highlighted), Add plugins and Web search. The Connectors submenu lists Add connector, Manage connectors and toggles for Descript, Gmail, Google Calendar and simonwillison.net. The Add connector submenu offers Browse connectors and Add custom connector.
8521
Simon Willison @simonwillison.net · 23/07/2026
Tucked away in this article is an appeal to the AI skeptics to PLEASE stop writing off stories like this OpenAI accidental exploit of Hugging Face as a dishonest marketing trick Frontier models can find and exploit vulnerabilities now, it helps nobody to pretend that they can't!
Resist the temptation to write this off as a stunt #

There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident.

To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!

The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that “autonomous exploit development by frontier AI agents is no longer a hypothetical capability”, and this incident is a perfect example of exactly that.
108910
Simon Willison @simonwillison.net · 22/07/2026
Today at Pier 39 in San Francisco I learned that if you want to take a good photograph of a sea lion, you should aim to get their teeth
A sea lion's head sticking out of the water showing off its two lower fangsAnother sea lion this time with its head only just showing, excellent whiskers, more visible lower teeth
4920