Sign in

Simon Willison

@simonwillison.net
51K followers 1.5K following 5K posts

Independent AI researcher, creator of datasette.io and llm.datasette.io, building open source tools for data journalism, writing about a lot of stuff at simonwillison.net

PostsRepliesMedia
Simon Willison @simonwillison.net · 6h
I included the reasoning traces in that one for some of the calculations, to show how the LLM handled long addition
Starting from the right:
2 + 0 = 2
2 + 7 = 9
6 + 9 = 15, write 5 carry 1
5 + 7 + 1 = 13, write 3 carry 1
0 + 6 + 1 = 7
1 + 0 = 1
6 + 4 = 10, write 0 carry 1
6 + 9 + 1 = 16, write 6 carry 1
3 + 7 + 1 = 11, write 1 carry 1
9 + 8 + 1 = 18, write 8 carry 1
9 + 0 + 1 = 10, write 0 carry 1
2 + 6 + 1 = 9
4 + 0 = 4

Wait, let me redo this more carefully.

4,299,366,105,622
6,088,794,067,970

Let me align them:
  4 2 9 9 3 6 6 1 0 5 6 2 2
  6 0 8 8 7 9 4 0 6 7 9 7 0

Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1
Position 4 (thousands): 5 + 7 + 1 = 13, write 3, carry 1
Position 5 (ten-thousands): 0 + 6 + 1 = 7
Position 6 (hundred-thousands): 1 + 0 = 1
Position 7 (millions): 6 + 4 = 10, write 0, carry 1
Position 8 (ten-millions): 6 + 9 + 1 = 16, write 6, carry 1
Position 9 (hundred-millions): 3 + 7 + 1 = 11, write 1, carry 1
Position 10 (billions): 9 + 8 + 1 = 18, write 8, carry 1
Position 11 (ten-billions): 9 + 0 + 1 = 10, write 0, carry 1
Position 12 (hundred-billions): 2 + 6 + 1 = 9
Position 13 (trillions): 4 + 0 = 4

Wait, that gives me 49,081,601,735,92? That doesn't seem right. Let me recount the digits.
2170
Simon Willison @simonwillison.net · 6h
I tried running it again with Qwen 3.8 27B on reasoning effort "medium" - just one run per calculation (not 30) because they took a lot longer It got them exactly right 167 out of 169 times, so the chart is pretty dull looking! Full details here: gist.github.com/simonw/8ef79...
Heatmap titled "Addition in words — Qwen3.8 27B — medium reasoning pilot", subtitle "1 fixed pair per ordered digit-length cell · easiest first (n = 169)". The x-axis is labeled "Number of digits in a" (1 to 13) and the y-axis is labeled "Number of digits in b" (1 to 13, bottom to top). A color bar labeled "Accuracy" runs from orange at 0.00 through white at 0.50 to blue at 1.00, with ticks at 0.00, 0.25, 0.50, 0.75 and 1.00. Every cell in the 13 by 13 grid reads 100% (blue) except two cells that read 0% (orange): a = 2 with b = 8, and a = 12 with b = 9.
1221
Simon Willison @simonwillison.net · 11h
Here are detailed (Astra-generated) notes on what I ran gist.github.com/simonw/88baf...
2120
Simon Willison @simonwillison.net · 16h
I was inspired to try re-running your experiment with a local, open weight model - Qwen 3.8 27B Q4A_K_M - here's the result
Heatmap titled "Addition in words — Qwen3.8 27B Q4_K_M", subtitle "30 fixed random pairs per ordered digit-length cell (n = 5,070)". The x-axis is labeled "Number of digits in a" (1 to 13) and the y-axis is labeled "Number of digits in b" (1 to 13, bottom to top). A color bar labeled "Accuracy" runs from red at 0.00 through yellow at 0.50 to green at 1.00, with ticks at 0.00, 0.25, 0.50, 0.75 and 1.00. Cell values for a = 1 to 13 in order, by row. b=1: 97%, 100%, 100%, 100%, 80%, 67%, 77%, 80%, 80%, 60%, 43%, 30%, 37%. b=2: 100%, 100%, 93%, 90%, 77%, 77%, 43%, 50%, 63%, 43%, 40%, 13%, 23%. b=3: 100%, 93%, 90%, 80%, 67%, 37%, 17%, 0%, 0%, 3%, 0%, 0%, 0%. b=4: 83%, 93%, 90%, 93%, 53%, 13%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b=5: 73%, 67%, 87%, 80%, 53%, 40%, 0%, 0%, 3%, 0%, 0%, 0%, 0%. b=6: 60%, 60%, 33%, 10%, 53%, 47%, 7%, 3%, 0%, 0%, 0%, 0%, 0%. b=7: 53%, 23%, 17%, 10%, 3%, 3%, 13%, 3%, 0%, 0%, 0%, 0%, 0%. b=8: 53%, 20%, 0%, 0%, 0%, 0%, 7%, 13%, 0%, 0%, 0%, 0%, 0%. b=9: 77%, 47%, 3%, 0%, 0%, 0%, 0%, 3%, 7%, 0%, 0%, 0%, 0%. b=10: 70%, 27%, 3%, 0%, 0%, 0%, 0%, 0%, 0%, 13%, 0%, 0%, 0%. b=11: 47%, 10%, then 0% for a = 3 to 13. b=12: 53%, 20%, then 0% for a = 3 to 13. b=13: 17%, 13%, then 0% for a = 3 to 13. Accuracy is high (green) in the bottom-left, where both numbers are short, and falls to 0% (red) across most of the upper-right.
5552
Simon Willison @simonwillison.net · 23h
The fact that it got the answers wrong a bunch of times suggest to me it didn't have a calculator, otherwise it surely would have got them all right
1320
Simon Willison @simonwillison.net · 29/09/2026
I'm at OpenAI's DevDay event in San Francisco today - as I have for the past three DevDay events, I'm running a live blog where I'll be posting updates during the keynote, which starts in five minutes simonwillison.net/2026/Sep/29/...
simonwillison.net
OpenAI DevDay 2026 live blog
I’m at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I’ll be live blogging the keynote and some other notes during the day. OpenAI gave me …
48811
Simon Willison @simonwillison.net · 29/09/2026
I wrote about why I think they've hit product market fit back in May simonwillison.net/2026/May/27/...
simonwillison.net
I think Anthropic and OpenAI have found product-market fit
Anthropic are strongly rumored to be about to have their first profitable quarter. Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by …
010
Simon Willison @simonwillison.net · 29/09/2026
Anthropic told investors that their Q2 was profitable, and it looks like their Q3 is going to be profitable too They haven't shared audited financials yet, I expect those will come with their S-1
030
Simon Willison @simonwillison.net · 29/09/2026
I don't think we have a Fable-class open weight model yet - wouldn't be surprised if we got there by the end of the year though
041
Simon Willison @simonwillison.net · 29/09/2026
Maybe if they can convince every other lab to slow down too!
250
Simon Willison @simonwillison.net · 29/09/2026
No idea, I think the fact that the Reuters story only includes 2025 numbers is really weird
190
Simon Willison @simonwillison.net · 29/09/2026
I'd love to know that We can at least estimate based on how much it costs to run similarly capable open weight models - Kiki K3 etc
170
Simon Willison @simonwillison.net · 29/09/2026
I'm pretty skeptical of those trillion dollar infrastructure numbers - they sound very bubbly to me Billions of dollars I can believe - Anthropic are already renting environmentally nasty data centers from SpaceX for over a billion dollars a month, and that's because they needed the extra capacity
050
Simon Willison @simonwillison.net · 29/09/2026
They're rumored to make a healthy margin on inference - so the more tokens they can sell the better Especially since corporate clients can't sign up for the $100 or $200/month subscription plans - companies have to pay the full API prices
3170
Simon Willison @simonwillison.net · 29/09/2026
They've caught up a bit in the past couple of months with Codex Desktop and GPT-5.6 and GPT-6 - but enterprise companies like signing longer deals and Anthropic snapped up a whole lot of those in the first half of the year
030
Simon Willison @simonwillison.net · 29/09/2026
OpenAI do seem a whole lot more enthusiastic than Anthropic at shouting about spending a trillion dollars on data centers or whatever - I think they blow a lot more hot air (which is saying a lot)
040
Simon Willison @simonwillison.net · 29/09/2026
OpenAI were blowing untold amounts of money on distractions like Sora - then in April they refocused because Anthropic were eating their lunch, and now they're trying to run the same playbook
240
Simon Willison @simonwillison.net · 29/09/2026
It was coding agents. Prior to Claude Code and Codex (and the Opus 4.5 / GPT 5.1 models that made those actually work) it was hard to spend more than $50/month on AI tokens, there just wasn't enough useful stuff to do with them A good coding agent can justify $50/day or more
5201
Simon Willison @simonwillison.net · 29/09/2026
Last year I didn't think OpenAI or Anthropic were long-term sustainable at all I thought their plan was to should "AGI! AGI!" at investors, keep the funding rolling in, and desperately hope to find a business model This year they found a business model, and it's not "AGI", it's coding agents
360
Simon Willison @simonwillison.net · 29/09/2026
The 2025 ones? Sure, if they kept on losing money at that rate in 2026. But all indications are that the opposite has happened - in 2026 they found product-market fit for a product that companies are spending thousands of dollars per employee per month on They didn't have that last year
3120
Simon Willison @simonwillison.net · 29/09/2026
Anthropic's only existed for five years, and the stories I've seen about them being financially sustainable have only been about 2026
190
Simon Willison @simonwillison.net · 29/09/2026
I mean Uber set a cap of $1,500 per employee per AI tool they were using, and if they have ~3,000 engineers on staff that's a cool $4.5m per month from just one customer, AFTER they set those limits
340
Simon Willison @simonwillison.net · 29/09/2026
Anthropic claimed that two of their quarters this year have been profitable - I'm holding out for the actual numbers in the full S-1 because of the many different ways they might be cooking the books, but I think there's a decent chance those claims will hold up
2120
Simon Willison @simonwillison.net · 29/09/2026
Yes, you're right - which makes their revenue acceleration in 2026 even more impressive
450
Simon Willison @simonwillison.net · 29/09/2026
And you know this! You know that a Reuters article about Anthropic's 2025 numbers is going to mislead people who are unaware of how much their revenue changed this year
2151
Simon Willison @simonwillison.net · 29/09/2026
Anthropic lost $8 billion in 2025 against $12 billion in revenue... but that was before they started pulling in that much revenue /per quatter/ in 2026 as companies started blowing millions of dollars on coding agent tools that hardly existed a year ago
3181
Simon Willison @simonwillison.net · 29/09/2026
Well yeah... your replies are filled with people who seem to think those 2025 numbers represent the current financial state of the company, when they very clearly do not
3220
Simon Willison @simonwillison.net · 28/09/2026
Hah I did actually think of you when I thought "who do I know called Mike", I figured a more common name was less risky! (I nearly went with "Tom")
110
Simon Willison @simonwillison.net · 28/09/2026
My hunch is that the challenge of regular non-nerds having trouble finding things to do with their agents will mostly be solved by word-of-mouth When your dumbass friend Mike figures out how to use Muse... Like how in the early days of MySpace people learned CSS to customize their page!
7320
Simon Willison @simonwillison.net · 28/09/2026
I really like how @reckless.bsky.social described this: "The people do not yearn for automation" www.theverge.com/podcast/9170...
theverge.com
BEWARE SOFTWARE BRAIN
Software brain is changing the world, but most people still aren’t buying.
5728
Simon Willison @simonwillison.net · 28/09/2026
Here's the video version on YouTube www.youtube.com/watch?v=GAkI...
youtube.com
WWC26-NA - 2026 in LLMs (so far)
YouTube video by WeAreDevelopers
0134
Simon Willison @simonwillison.net · 28/09/2026
(The full talk also covers New Zealand’s Kākāpō parrot breeding season, probably the best news of 2026)
1100
Simon Willison @simonwillison.net · 28/09/2026
And my closing thought on how this stuff impacts our lives as software engineers, with a quote from cycling champion Greg LeMond
It doesn't get easier -
you just get faster
Greg LeMond
3x Tour de France champion
120
Simon Willison @simonwillison.net · 28/09/2026
Here's how I define "Fable class models" - first Claude Fable 5, now Claude Opus 5.5 and GPT-Astra 6 and maybe GPT-5.6 Sol as well
Fable class models
If you can define a goal, provide unambiguous instructions,
and provide access to necessary tools They can solve your
problem with brute force
240
Simon Willison @simonwillison.net · 28/09/2026
I've published detailed notes and an annotated transcript to accompany the video of the keynote I gave at @wearedevelopers.bsky.social World Congress North America in San Jose on Friday - here's my rundown of everything that's happened with LLMs in 2026 so far simonwillison.net/2026/Sep/27/...
simonwillison.net
2026 in LLMs (so far)
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …
310420
Simon Willison @simonwillison.net · 27/09/2026
She made nesting bowls too! bsky.app/profile/simo...
030
Simon Willison @simonwillison.net · 27/09/2026
JPEG XL looks not to have good browser support yet, but it looks like AVIF became usable around 2023 caniuse.com?search=AVIF
Screenshot of the CanIUse support charts for AVIF - Chrome, Firefox, Opera have had support since before 2022, Safari and Safari on iOS gained support mid-2022, Edge finally got there at the end of 2023
120
Simon Willison @simonwillison.net · 27/09/2026
I just exported 69 slides from Keynote - as (reduced quality) JPG: 10.4MB, as PNG: 100MB, as WebP: 4.1MB Quality looks great, this 1920×1080 image here is 58KB
6260
Simon Willison @simonwillison.net · 27/09/2026
How late to the party am I on WebP? I only recently started using it as my preferred format for non-photo images and it's SO MUCH smaller than PNG, and looks much better than JPG used to shrink down screenshots too Has everyone else been having a WebP party without me for years?
16751
Simon Willison @simonwillison.net · 27/09/2026
But really the main thing is to poke around with them, try to get them to build impossible things, and see what they still can't do yet
120
Simon Willison @simonwillison.net · 27/09/2026
Feels a bit self-serving to say it, but my blog has a LOT of material on coding agents simonwillison.net/tags/coding-... - plus an in-progress guide that I need to invest more work in simonwillison.net/guides/agent...
simonwillison.net
Simon Willison on coding-agents
252 posts tagged ‘coding-agents’. Systems where an LLM writes code which is then compiled, executed, tested or otherwise exercised by tools in a loop.
250
Simon Willison @simonwillison.net · 27/09/2026
Plus saying "have you been following this year's Kākāpō breeding season?" is a fantastic conversation-starter
1130
Simon Willison @simonwillison.net · 27/09/2026
I am now the owner of some VERY fine Kākāpō ceramics, and the lego set, and I have a plushy somewhere too bsky.app/profile/simo...
2180
Simon Willison @simonwillison.net · 27/09/2026
Sirocco on Last Chance To See, followed by stumbling across the Kākāpō Recovery Programme first on Twitter and now on Bluesky, in particular @digs.bsky.social
230
Simon Willison @simonwillison.net · 27/09/2026
Prompt and transcripts here simonwillison.net/2026/Sep/26/... - or you can load the HTML page Claude built to see the interactive version before it became a video tools.simonwillison.net/kakapo-party
tools.simonwillison.net
Kākāpō Party
1160
Simon Willison @simonwillison.net · 27/09/2026
I included a few references to this year's record-breaking Kākāpō breeding season in a talk I gave yesterday, and since Claude Opus 5.5 is surprisingly capable at pixel art animation I had it create this celebratory video for my closing slide
6835
Simon Willison @simonwillison.net · 24/09/2026
I've seen less since the Pacifica Pier got popular, I think we have lost a few to the new location
020
Simon Willison @simonwillison.net · 23/09/2026
Seems to cost less than 1 cent per minute of generated audio for Flash, and Flash-Lite is even cheaper than that
2170
Simon Willison @simonwillison.net · 23/09/2026
The new Gemini 3.8 TTS models are super-cheap and can generate conversations between multiple voices (from 2,000+, or you can clone your own) - I built a little playground UI for it, then had Claude knock up a script where two pelicans debate moving to Pacifica Pier simonwillison.net/2026/Sep/23/...
9847
Simon Willison @simonwillison.net · 23/09/2026
I'm not using Max in Claude Code
100