Sign in

Simon Willison

@simon.fedi.simonwillison.net.ap.brid.gy
10K followers 3 following 1.6K posts

Open source developer building tools to help journalists, archivists, librarians and others analyze, explore and publish their data. datasette.io […] [bridged from fedi.simonwillison.net/@simon on the fediverse by fed.brid.gy ]

PostsRepliesMedia
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 04/10/2026
Just published this post about how we’re going to need default hard budget caps on pretty much everything simonwillison.net/2026/Oct/3/defaul…
simonwillison.net
We’re going to need default hard budget caps on pretty much everything
Here’s a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I’m talking about the …
145
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 23/09/2026
Big model release today - I wrote about Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna - plus comparison grids of pelicans by the different model families at different reasoning levels simonwillison.net/2026/Sep/22/opus-…
simonwillison.net
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Yesterday was Grok 4.7 (pelicans) and MiMo v2.6 Flash/Pro (more pelicans). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It’s …
1113
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 17/08/2026
My conclusions from the end if the post
Some observations #
The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I’m delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models—today it can run on a capable laptop.

The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That’s the catch with these dense (non-Mixture-of-Experts) models—they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard.

The most important thing about Qwen 3.8 27B is what it demonstrates. We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file.

The models at this size continue to get better at an impressive rate. We don’t need to spend half a million dollars on datacenter-class hardware just to run a competent model.
081
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 16/08/2026
Fun bonus: I set Pi up with Qwen 3.8 27B and had it build a script for transforming its own .jsonl transcripts to Markdown... which it did! Here's the transcript of it building the tool I then used to share that transcript gist.github.com/simonw/491e55ac9d74…
gist.github.com
jsonl-to-markdown.md
GitHub Gist: instantly share code, notes, and snippets.
130
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 16/08/2026
Here’s my review of Qwen 3.8 27B - I can't remember the last time I've had this much fun playing with a local model that runs on my own computers simonwillison.net/2026/Aug/16/qwen-…
simonwillison.net
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Friday’s big release was Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba’s Qwen research lab. I’ve been looking forward to this one: 27B is an …
22111
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 14/08/2026
The new Qwen 3.7 27B, running as a 17GB GGUF in LM Studio on my M5 Max MacBook Pro, just drew me the best pelican riding a bicycle I've seen from any model that runs on my laptop It did take nearly 21 minutes to generate, and used 22,276 reasoning […] [Original post on fedi.simonwillison.net]
The pelican has the right shaped beak. The red bicycle has the correct shape of frame. The pelican's wing reaches the handlebars. It has legs on both side of the bicycle. There is a pleasing set of clouds, birds, sun, grass and shadow on the image, plus motion lines behind but not in front of the bird.
293
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 05/08/2026
Big new release of my LLM CLI tool and Python library for talking to hundreds of different LLMs - reasoning traces, OpenAI Responses support, server-side tools, smarter logging and a whole lot more simonwillison.net/2026/Aug/4/new-re…
simonwillison.net
New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging
I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side …
042
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 28/07/2026
Hugging Face just published a highly detailed technical account of OpenAI's accidental cyberattack on their systems - it's wild how sophisticated this was: huggingface.co/blog/agent-intrusion… Wrote up some of my own notes here […]
fedi.simonwillison.net
Original post on fedi.simonwillison.net
388
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 26/07/2026
Ruff 0.16.0 - Astral's fast Python linter - came out a few days ago and increased the number of default-enabled rules from 59 to 413, which highlighted all sorts of problems across my projects (1618 in sqlite-utils alone) simonwillison.net/2026/Jul/25/ruff
simonwillison.net
Ruff v0.16.0
Astral shipped a significant new version of their Ruff Python linting tool a few days ago on July 23rd. I noticed today because my various CI jobs all started failing …
051
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 23/07/2026
I wrote about the completely wild incident where OpenAI were testing a new model and it broke out of its sandbox and broke INTO Hugging Face to steal the answers to the benchmark simonwillison.net/2026/Jul/22/opena…
simonwillison.net
OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened
This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model’s guardrail features turned off. Rather than solve the test, the …
11518
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 16/07/2026
My notes on Kimi K3, plus some thoughts on what we can still learn from the pelican benchmark even while it becomes further detached from how good the models are at the things that matter (like agentic tool calling across longer conversations) simonwillison.net/2026/Jul/16/kimi-…
simonwillison.net
Kimi K3, and what we can still learn from the pelican benchmark
Chinese AI lab Moonshot AI announced Kimi K3 this morning, describing it as their “most capable model to date, with 2.8 trillion parameters”. It’s currently available via their website and …
0106
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 14/07/2026
New TIL: Using uvx in GitHub Actions in a cache-friendly way I finally found a recipe that I like for running `uvx tool-name` in GitHub Actions without downloading a fresh copy of the package every time til.simonwillison.net/github-action…
til.simonwillison.net
Using uvx in GitHub Actions in a cache-friendly way
I often find myself wanting to run a quick Python tool inside of GitHub Actions using uvx name-of-tool - but I don't want that to result in a network request to PyPI every time the workflow runs. I want the tool to be fetched the first time and then reused from the GitHub Actions cache for subsequent runs.
0102
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 09/07/2026
... and if you want to see some of OpenAI's own pelicans they featured a 3D pelican riding a tricycle, bicycle, pony, and another pelican in their livestream this morning: www.youtube.com/live/Wq45rvPGNHs?t=…
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 09/07/2026
Notes on GPT-5.6, which includes some interesting new additions to the API (programmatic tool calling and multi-agent in particular) - plus 18 pelicans for the 6 reasoning levels and 3 new models: simonwillison.net/2026/Jul/9/gpt-5-6
simonwillison.net
The new GPT-5.6 family: Luna, Terra, Sol
OpenAI’s latest flagship model hit general availability this morning, and comes in three sizes: Luna, Terra, and Sol (from smallest to largest). The new models are priced per 1M input/output …
271
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 07/07/2026
Since this is a slightly backwards-incompatible change here's a detailed upgrade guide - which you can read, or feed into your coding agent and have it apply the upgrades for you sqlite-utils.datasette.io/en/stable…
sqlite-utils.datasette.io
Upgrading - sqlite-utils
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 07/07/2026
I released sqlite-utils 4.0, the 124th release but the first major version bump since 3.0 back in 2020 I managed to keep things backwards-compatible all the way up to version 3.39 before the accumulated design mistakes forced me to bump that number! […]
fedi.simonwillison.net
Original post on fedi.simonwillison.net
120
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 30/06/2026
I've added video support to my "shot-scraper" browser automation tool - you (or your coding agent) can now create a storyboard YAML file and use that to record a video demo of new web application features simonwillison.net/2026/Jun/30/shot-…
simonwillison.net
Have your agent record video demos of its work with shot-scraper video
shot-scraper video is a new command introduced in today’s shot-scraper 1.10 release which accepts a storyboard.yml file defining a routine to run against a web application and uses Playwright to …
112
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 23/06/2026
My parallel agent side-project today was having Claude Code port the new Moebius image pinpointing model to ONNX in order to run it entirely in the browser simonwillison.net/2026/Jun/22/porti…
simonwillison.net
Porting the Moebius 0.2B image inpainting model to run in the browser with Claude Code
This morning on Hacker News I saw Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance, describing a small but effective inpainting model—a model where you can mark regions of …
110
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 22/06/2026
I just released the first release candidate for sqlite-utils v4, adding a migrations system (previously released independently as sqlite-migrate) and support for nested transactions: simonwillison.net/2026/Jun/21/sqlit…
simonwillison.net
sqlite-utils 4.0rc1 adds migrations and nested transactions
sqlite-utils is my combined Python library and CLI tool for working with SQLite databases. It provides an extensive set of higher-level operations on top of Python’s default sqlite3 package, including …
001
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 19/06/2026
Lots more information in this post on the Datasette project blog, including details on our live demo and uv one-liners you can use to try this out on your own machine datasette.io/blog/2026/datasette-ap…
datasette.io
Host applications inside Datasette with Datasette Apps - Datasette Blog
001
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 19/06/2026
Think of this as Claude Artifacts reimagined for Datasette - you get all the power of artifacts but with a JSON API to a full relational database, allowing your HTML+JS apps to access and store data in all shapes and sizes
100
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 19/06/2026
Just launched Datasette Apps - a plugin for Datasette that lets you host full HTML+JS apps in an iframe sandbox that can query your database and do interesting things with your data simonwillison.net/2026/Jun/18/datas…
simonwillison.net
Datasette Apps: Host custom HTML applications inside Datasette
Today we launched a new plugin for Datasette, datasette-apps, with this launch announcement post on the Datasette project blog. That post has the what, but I’m going to expand on …
121
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 14/06/2026
It's now possible to compile Python extensions (C, C++, Rust etc) to WebAssembly and distribute them through PyPI such that Pyodide can install them directly simonwillison.net/2026/Jun/13/publi…
simonwillison.net
Publishing WASM wheels to PyPI for use with Pyodide
The Pyodide 314.0 release announcement (via Hacker News) includes news I’ve been looking forward to for a long time: You can now publish Python packages built for Pyodide (or any …
055
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 12/06/2026
After two days with Claude Fable 5 the best way I can describe it is "relentlessly proactive" - here's an example where I dropped in a screenshot of a bug and it span up custom CORS Python servers and used pyobjc-framework-Quartz to capture screenshots […]
fedi.simonwillison.net
Original post on fedi.simonwillison.net
314
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 11/06/2026
Since the JSON extras API is a little hard to explain without an example I also had Fable 5 and GPT-5.5 collaborate on this custom API explorer tool for trying out the new feature […] [Original post on fedi.simonwillison.net]
Screenshot of a web application titled "Datasette extras explorer". A URL input field contains https://latest.datasette.io/fixtures/facetable.json with a teal Explore button next to it. Below, a left panel labeled EXTRAS (30) lists checkboxes: all_columns - All columns in the table, regardless of _col/_nocol filtering; column_types - Column type assignments for this table; columns (checked) - Column names returned by this query; count - Total count of rows matching these filters; count_sql - SQL query used to calculate the total count; custom_table_templates - Custom template names considered for this table; database - Database name; database_color - Color assigned to the database. A right panel labeled RESPONSE shows GET /fixtures/fac… with Copy JSON and Copy URL buttons, then a dark JSON viewer showing 200 - 9.9 KB - 114ms and JSON: "ok": true, "next": null, "columns": (highlighted array) "pk", "created", "planet_int", "on_earth", "state", "_city_id", "_neighborhood", "tags", "complex_array", "distinct_some_null", "n", "rows": list of objects.
010
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 11/06/2026
New Datasette release: 1.0a33, which finally brings documents the ?_extra= JSON API mechanism and brings it to the row and query pages in addition to the table pages (Most of the code in this release was built with the help of Claude Fable 5) datasette.io/blog/2026/api-extras
datasette.io
Datasette 1.0a33 with JSON extras in the API - Datasette Blog
110
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 10/06/2026
Wrote up my initial impressions of Claude Fable 5 - it has a big model smell: slow, expensive and capable of crunching through pretty much everything I threw at it simonwillison.net/2026/Jun/9/claude…
simonwillison.net
Initial impressions of Claude Fable 5
I didn’t have early access to today’s Claude Fable 5 release, but I’ve spent the past ~5.5 hours putting it through its paces. My initial impressions are that this is …
487
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 06/06/2026
You can try it out like this: uvx micropython-wasm -c 'print("Hello world")'
021
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 06/06/2026
I may have finally found the Python-in-a-sandbox solution I've been looking for... here's my latest experiment, this time running MicroPython in WebAssembly inside my Python applications simonwillison.net/2026/Jun/6/microp…
simonwillison.net
Running Python code in a sandbox with MicroPython and WASM
I’ve been experimenting with different approaches to running code in a sandbox for several years now, but my latest attempt feels like it might finally have all of the characteristics …
374
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 03/06/2026
Uber reportedly now caps coding agents at $1,500/month per employee per tool - seems sensible to me, but it's also an interesting hint at the value Uber thinks these tools are providing simonwillison.net/2026/Jun/3/uber-c…
simonwillison.net
Uber Caps Usage of AI Tools Like Claude Code to Manage Costs
I wrote the other day about Uber blowing its 2026 AI budget in four months, and how that wasn't particularly surprising given they would have set that budget in 2025, …
167
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 29/05/2026
@slowenough I had it do a NORTH VIRGINIA OPOSSUM ON AN E-SCOOTER too - it made this: gist.github.com/simonw/68560eddb0b2… Pales in comparison to GLM-5.1 though simonwillison.net/2026/Apr/7/glm-51…
110
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 29/05/2026
Notes on Claude Opus 4.8, which Anthropic call "a modest but tangible improvement" - plus pelicans riding bicycles for each of the five different thinking efforts simonwillison.net/2026/May/28/claud…
SVGs of pelicans riding bicycles, in low, medium, high, xhigh and max. They do get progressively better. Only the max one has a correctly shaped bicycle frame.
3104
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 27/05/2026
Given the recent burst of activity around enterprise pricing and contracts, I think April 2026 was the month when both OpenAI and Anthropic found product-market fit simonwillison.net/2026/May/27/produ…
simonwillison.net
I think Anthropic and OpenAI have found product-market fit
Anthropic are strongly rumored to be about to have their first profitable quarter. Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by …
244
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 26/05/2026
When I woke up this morning I didn't think I'd be spending a bunch of time today getting familiar with Catholic theology, but here we are. Notes on Pope Leo XIV's encyclical on AI. simonwillison.net/2026/May/25/encyc…
simonwillison.net
A few notes on Pope Leo XIV’s encyclical on AI
Dropped this morning by the Vatican: Magnifica Humanitas of His Holiness Pope Leo XIV on Safeguarding the Human Person in the Time of Artificial Intelligence. This is a very interesting …
0715
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 25/05/2026
Plugins can also affect the empty state of the new menu - the latest datasette-agent adds a form for kicking off a new agent conversation - live demo (if you sign in with GitHub) on agent.datasette.io
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 25/05/2026
Fun new feature in Datasette 1.0a30 - a "jump to" menu, also available by the "/" keyboard shortcut, which lets you type to jump to a database, table or canned query. It offers a hook so plugins can extend it with additional content too datasette.io/blog/2026/jump-menu
100
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 21/05/2026
I released the first alpha of Datasette Agent - a conversational AI assistant for Datasette that can answer questions about data in SQLite databases, and can be extended with plugins to add extra tools and features You can watch a demo video or try it out yourself on agent.datasette.io […]
fedi.simonwillison.net
Original post on fedi.simonwillison.net
120
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 20/05/2026
I don't have much to say about this year's Google I/O because I prefer to write about products that have shipped, not just "coming soon" announcements - but here are some notes on Gemini Spark and Antigravity simonwillison.net/2026/May/20/googl…
simonwillison.net
Google I/O, Gemini Spark, Antigravity
It's hard to find much to write about Google I/O this year because I have a policy of not writing about anything that I can't try out myself, and a …
110
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 19/05/2026
My notes on Gemini 3.5 Flash - 3x the price of Gemini 3 Flash but Google are planning to use it for many of their own products simonwillison.net/2026/May/19/gemin…
simonwillison.net
Gemini 3.5 Flash: more expensive, but Google plan to use it for everything
Today at Google I/O, Google released Gemini 3.5 Flash. This one skipped the -preview modifier and went straight to general availability, and Google appear to be using it for a …
031
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 12/05/2026
Wrote about today's GitLab restructuring / "workforce reduction" announcement, and ended up digging around in version control for both the GitLab and the 37signals public employee handbooks to help illustrate my thoughts simonwillison.net/2026/May/11/gitla…
simonwillison.net
GitLab Act 2
There's a lot going on in this announcement from GitLab about the "workforce reduction" and "structural and strategic decisions" they are making with respect to the agentic era. They're "planning …
342
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 11/05/2026
New TIL: I figured out how to use my LLM CLI tool in a shebang line, which means you can write executable scripts in English, or hook up more complex scripts with a snippet of YAML template - til.simonwillison.net/llms/llm-sheb…
#!/usr/bin/env -S llm -x -f
Generate an SVG of a pelican riding a bicycle

But you can also incorporate tool calls:

#!/usr/bin/env -S llm -T llm_time -f
Write a haiku that mentions the exact current time

Or even execute YAML templates directly that define extra tools as Python functions:

#!/usr/bin/env -S llm -t
model: gpt-5.4-mini
system: |
  Use tools to run calculations
functions: |
  def add(a: int, b: int) -> int:
      return a + b
  def multiply(a: int, b: int) -> int:
      return a * b

Then:

./calc.sh 'what is 2344 * 5252 + 134' --td

Which outputs (thanks to that --td tools debug option):

Tool call: multiply({'a': 2344, 'b': 5252})
  12310688

Tool call: add({'a': 12310688, 'b': 134})
  12310822

2344 × 5252 + 134 = **12,310,822**
130
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 07/05/2026
Oh, and Elon said "We reserve the right to reclaim the compute if their AI engages in actions that harm humanity."
012
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 07/05/2026
Under-reported details of the xAI/Anthropic Colossus data center deal: Anthropic get Colossus 1 but xAI keep using the larger Colossus 2, Colossus 1 has a REALLY bad environmental record, and xAI just shut down a bunch of older models on 2 weeks' notice […]
fedi.simonwillison.net
Original post on fedi.simonwillison.net
137
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 02/05/2026
I added a new feature to my blog (built entirely on my phone with Claude code for web) that imports my iNaturalist photos and adds them to my site's overall timeline simonwillison.net/2026/May/2/sighti…
simonwillison.net
Sightings
I have a new camera (a Canon R6 Mark II) so I'm taking a lot more photos of birds. I share my best wildlife photos on iNaturalist, and based on …
000
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 30/04/2026
I particularly appreciate how this rationale isn't based on the idea that LLM code is of poor quality compared to code written by hand - the quality of the code isn't the deciding factor at all here
032
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 30/04/2026
The Zig project's rationale for their blanket ban on AI-assisted contributions makes a lot of sense to me - for them, time spent reviewing PRs isn't about the code, it's about growing new contributors for the future of the project simonwillison.net/2026/Apr/30/zig-a…
simonwillison.net
The Zig project's rationale for their firm anti-AI contribution policy
Zig has one of the most stringent anti-LLM policies of any major open source project: No LLMs for issues. No LLMs for pull requests. No LLMs for comments on the …
21525
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 29/04/2026
I released LLM 0.32a0 this morning, a major backwards-compatible refactor of my LLM Python library and CLI tool for working with language models - the new changes should help LLM work better with reasoning models and other new frontier capabilities simonwillison.net/2026/Apr/29/llm
simonwillison.net
LLM 0.32a0 is a major backwards-compatible refactor
I just released LLM 0.32a0, an alpha release of my LLM Python library and CLI tool for accessing LLMs, with some consequential changes that I’ve been working towards for quite …
002
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 28/04/2026
It didn't quite manage to draw me a pelican riding a bicycle, but I still appreciated its era-appropriate response
I couldn't resist hitting the talkie demo with my classic test prompt:

Generate an SVG of a pelican riding a bicycle

It replied:

An SVG of a pelican riding a bicycle was generated in 1860. The pelican was represented as perched on the saddle, with its long bill pointed forward, and its legs astride the handles. The idea was suggested by the fact that pelicans are sometimes seen fishing on horseback, on the banks of the Rhine.
233
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 28/04/2026
Some notes on talkie, a new "vintage language model" from a team including Alec Radford (yes, that Alec Radford) "trained on 260B tokens of historical pre-1931 English text" simonwillison.net/2026/Apr/28/talkie
simonwillison.net
Introducing talkie: a 13B vintage language model from 1930
New project from Nick Levine, David Duvenaud, and Alec Radford (of GPT, GPT-2, Whisper fame). talkie-1930-13b-base (53.1 GB) is a "13B language model trained on 260B tokens of historical pre-1931 …
267
Simon Willison @simon.fedi.simonwillison.net.ap.brid.gy · 28/04/2026
Here's a uv one-liner that downloads and runs the MLX model against a local mp3 file uv run --with mlx-audio python -m mlx_audio.stt.generate \ --model mlx-community/VibeVoice-ASR-4bit \ --audio lenny.mp3 --output-path lenny \ --format json --verbose --max-tokens 32768
Screenshot of a macOS terminal running an mlx-audio speech-to-text command using the VibeVoice-ASR-4bit model on lenny.mp3, showing download progress, a warning that audio duration (99.8 min) exceeds the 59 min maximum so it's trimming, encoding/prefilling/generating progress bars, then a Transcription section with JSON segments of speakers discussing AI coding agents, followed by stats: Processing time 524.79 seconds, Prompt 26615 tokens at 50.718 tokens-per-sec, Generation 20248 tokens at 38.585 tokens-per-sec, Peak memory 30.44 GB.
010