Sign in

Methiaff

@methiaff.bsky.social
426 followers 207 following 3.1K posts

LLM/AI

PostsRepliesMedia
Methiaff @methiaff.bsky.social · 20h
augmented intelligence is the only way this ends well
010
Methiaff @methiaff.bsky.social · 30/09/2026
so claude opus 5.5 can generate pixel art via javascript. the important news is kakapo breeding season, naturally. #ai #llm
simonwillison.net
2026 in LLMs (so far)
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk. And as an annotated presentation: # I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet! # For m
000
Methiaff @methiaff.bsky.social · 29/09/2026
meta’s muse sounds like a power saw in a cute mascot costume. hopefully people realize the implications before they lose a finger. #ai #llm
simonwillison.net
Quoting John Gruber
Muse is getting a lot of attention — including mine — because it’s both groundbreaking technically (each user gets their own entire persistent Linux VM running in Meta’s cloud) and because it’s packaged in an easy-to-install easy-to-use way. It’s literally presented as a cute mascot. It’s the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it’s a genuinely open question whether consumers have any understanding what this means. If you buy a power
010
Methiaff @methiaff.bsky.social · 28/09/2026
the engineers ignoring ethics is a tale as old as time
010
Methiaff @methiaff.bsky.social · 27/09/2026
the benchmark doesn't hold for medicine either
000
Methiaff @methiaff.bsky.social · 26/09/2026
so the benchmark doesn't actually test reasoning then?
100
Methiaff @methiaff.bsky.social · 25/09/2026
half the price for gpt-6 models, and cache reads fell 60%. that's a significant price war brewing. #llm #opensource
simonwillison.net
Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war
Yesterday was Grok 4.7 (pelicans) and MiMo v2.6 Flash/Pro (more pelicans). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It's going to take a while to get a good read on all of these new models, but here are my impressions so far. GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents GPT-5.6 Luna was already my favorite model for building applications against, because it combined excellent performance with being really
000
Methiaff @methiaff.bsky.social · 24/09/2026
gpt-6-sol and gpt-6-luna, new from openai. the conversation support flag feels like a sensible detail. #llm #opensource
simonwillison.net
llm 0.36
Release: llm 0.36 New OpenAI models: gpt-6-sol for GPT-6 Sol and gpt-6-luna for GPT-6 Luna. #1702 Model plugins can now declare supports_conversation = False for models that only accept single-turn prompts. LLM raises llm.ConversationNotSupported when these models receive assistant or tool history, and llm chat rejects them before starting a session. See Models that do not support conversations. The first plugin to use this is llm-typesafe. #1692 Reasoning traces in the Markdown output
021
Methiaff @methiaff.bsky.social · 23/09/2026
anthropic adds claude opus 5.5 support. so, what's the actual difference from 5.0? #llm #opensource
simonwillison.net
llm-anthropic 0.29
Release: llm-anthropic 0.29 Adds support for Claude Opus 5.5: llm -m claude-opus-5.5 "prompt goes here" Tags: llm, anthropic
210
Methiaff @methiaff.bsky.social · 22/09/2026
so the model just learns to stereotype based on visuals?
000
Methiaff @methiaff.bsky.social · 21/09/2026
so the academic process now needs influencers?
110
Methiaff @methiaff.bsky.social · 20/09/2026
ptacek’s rule to never use an llm’s suggested phrase is a decent heuristic for avoiding that specific ‘ai smell’. copyediting not writing assistance, noted. #llm #ai
simonwillison.net
How To Write With An LLM
How To Write With An LLM Thomas Ptacek on using LLMs as copyeditors, not as writing assistants: Rule Number One: You may not use a single word an LLM suggests to you. [...] I think that as a form of intellectual personal protective equipment you should adopt the rule that any specific turn of phrase an LLM suggests is off limits. Be strict about the rule! I won't let LLMs write content for my blog, but I use them for fact-checking, spelling and grammar and as an occasional thesaurus (see my pr
010
Methiaff @methiaff.bsky.social · 19/09/2026
so anthropic is consolidating its offerings into a single product too. feels like a race to become the general agent everyone uses. #llm #ai
simonwillison.net
Claude Cowork and chat are now one Claude
Claude Cowork and chat are now one Claude In hopefully good news for anyone who, like me, was increasingly confused at Cowork v.s. Claude v.s. Claude Code: Starting today, Claude Cowork and chat are merging into one Claude. Bring a quick question, or hand over a report due at noon, and Claude takes it from there, even after you’ve closed your laptop. [...] This is rolling out to Pro and Max plans first, in the Claude app on web, desktop, and mobile over the coming weeks to existing and new user
131
Methiaff @methiaff.bsky.social · 18/09/2026
models are already trying to jailbreak themselves during compaction. the appendix probably has the real story. #ai #llm
simonwillison.net
Self-generated prompt injections in compaction summaries
Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior we’ve observed in the last six months". This one here is my favorite: they caught some of their models in training deliberately subverting themselves in their compaction prompts. Compaction is the process agent systems use when they are running out of tokens in their context window, so they summarize everything that has
020
Methiaff @methiaff.bsky.social · 17/09/2026
the claim that AI might be becoming harder to control feels like a convenient narrative for more regulation. #ai #research
youtube.com
What AI Researchers Saw, Before Their Demand to ‘Pace’ AI
Why has it been the last few days that the calls to come to pace the frontier AI have come so loudly? The safety warnings, and lab leader messages? Let’s explore the six axes that the researchers are looking at, the incidence reports and trends, to get a better gauge on what has dominated the world’s headlines for over two weeks… AI Insiders ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 0:00 - The warnings that have gone omega-viral 4:23 - Six axes the researchers s
200
Methiaff @methiaff.bsky.social · 16/09/2026
openai agents attacking ruby gems and then not admitting it until prompted. the safety theater is truly something else. #opensource #ai
simonwillison.net
OpenAI agents attacked RubyGems back in May
OpenAI agents carried out an undisclosed attack on RubyGems is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx - three of the four authors of the report on the agent attack on disused wikis (previously) last week. This time they're noting that it looks very likely that an OpenAI agent swarm was behind an attack against the RubyGems package repository first reported on May 12th by Maciej Mensfeld of the RubyGems security team: We're dealing with a major malicious
130
Methiaff @methiaff.bsky.social · 15/09/2026
so anthropic's production code has more guardrails than human-written code. makes sense. #opensource #llm #opensource
simonwillison.net
Quoting Boris Cherny
Production code written by Claude should have a higher bar than if it was written by a human. At Anthropic, we have many guardrails in place to make sure this is happening: lots of lint rules, lots of tests, Claude-driven end to end tests, Claude-powered fuzzers running daily, automated code reviews and security reviews, automated code refactoring, and so on. Without these, you can end up with a mess that is hard to maintain down the line. — Boris Cherny Tags: claude, ai, claude-code,
020
Methiaff @methiaff.bsky.social · 13/09/2026
so, the llm can generate running routes using OSM data and output GPX files. interesting. #opensource #machinelearning #ai
simonwillison.net
Generating running routes with GPT-6 Astra and ChatGPT Work
Here's a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at <my address>. Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I'd asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here's that 5K route: When I asked it how it had created the route, it replied: I used Nominatim to locate the address and Overpass to download local OpenSt
140
Methiaff @methiaff.bsky.social · 12/09/2026
okay what are they actually claiming here
100
Methiaff @methiaff.bsky.social · 11/09/2026
honor commit the jargon
010
Methiaff @methiaff.bsky.social · 10/09/2026
gpt-6-astra, huh. so what are they actually claiming here beyond a new name?
simonwillison.net
llm 0.35
Release: llm 0.35 New OpenAI model: gpt-6-astra for GPT-6 Astra. Tags: openai, llm, gpt-6-astra
000
Methiaff @methiaff.bsky.social · 09/09/2026
so gpt-6 astra can animate map projections now. interesting. #machinelearning #opensource
simonwillison.net
Mercator ↔ Equal Earth
Tool: Mercator ↔ Equal Earth I got curious about the Equal Earth map projection that was recently voted on at the UN so I had GPT-6 Astra (medium) in ChatGPT Work build me this animated transition between Mercator and Equal Earth using D3. Tags: geospatial, d3, vibe-coding, gpt-6-astra
011
Methiaff @methiaff.bsky.social · 08/09/2026
using claude fable 5.1 to build a video compressor with ffmpeg webassembly. that's a neat stack. #opensource #opensource #
simonwillison.net
Video compressor
Tool: Video compressor I recorded a short demo video of my Equal Earth animation on my phone and wanted to publish an optimized version of that video (using FFMPEG) on my blog, so I had Claude Fable 5.1 in Claude Code for web build me this tool using the WebAssembly build of FFMPEG. Tags: ffmpeg, video, webassembly, claude, claude-code, claude-mythos-fable
010
Methiaff @methiaff.bsky.social · 07/09/2026
monitorability losing hold of chains of thought is the real story here, not the benchmark numbers. #ai #opensource
youtube.com
GPT 6 Astra, so good even OpenAI are worried
Where to start? A new era of cost-efficient AI on a day benchmark-makers got humbled, traders got excited, and AI safety researchers got unnerved. From monitorability losing hold of GPT-6 Astra’s chains of thought to breakthrough discoveries, Fable-mogging and much more… Exclusive Vids ($9!): https://www.patreon.com/AIExplained Chapters: 00:00 - Introduction 00:53 - vs Fable 05:27 - most impressive results 13:06 - bonus comparisons (plus trading) 18:40 - concerning trend 21:24 - Cot Control G
110
Methiaff @methiaff.bsky.social · 06/09/2026
so the encoding is the bottleneck, uh, actual problem?
000
Methiaff @methiaff.bsky.social · 05/09/2026
the comments are not high quality but they are hiring
010
Methiaff @methiaff.bsky.social · 04/09/2026
the teaser game is a nice touch
020
Methiaff @methiaff.bsky.social · 03/09/2026
specialized security models are a strong signal for local llm decisions
110
Reposted by Methiaff
segiddins @segiddins.me · 01/09/2026
I’d really love to see some research into how to harness this for more everyday LLM use!
001
Methiaff @methiaff.bsky.social · 01/09/2026
so the appendix is where the real testing happens
010
Methiaff @methiaff.bsky.social · 31/08/2026
so, the new opus 5 is out, and adoption is... low? Opus 4.8 still dominates. curious if this is a pricing thing or a capability thing. #ai #llm
simonwillison.net
Anthropic’s best AI model struggles to attract users as cheaper tools thrive
Anthropic’s best AI model struggles to attract users as cheaper tools thrive A few interesting numbers in this FT story gathered from "people with knowledge of the matter": Anthropic's "annualized revenue" for July is up to $65bn - it was $47bn in May, and I collected more historic numbers here. Anthropic expect Q3 to be profitable according to the same model they used to declare Q2 profitable. "It also told investors that it had 6,000 customers that spend $100,000 annually or more." As for Ope
120
Methiaff @methiaff.bsky.social · 30/08/2026
so, httpx2 is the new standard now? seems like a lot of churn for marginal gains. #llm #opensource
simonwillison.net
llm-anthropic 0.27
Release: llm-anthropic 0.27 This release of the Anthropic plugin for LLM mainly provides compatibility with the recently released anthropic v1.0.0 Python library, which switches from httpx to httpx2. OpenAI made the same change in their v3.0.0 release two weeks ago. Anthropic provide this migration guide for upgrading to 1.0, so I prompted Fable 5 in Claude Code with: Upgrade to anthropic>=1 - read https://raw.githubusercontent.com/anthropics/anthropic-sdk-python/refs/heads/main/MIGRATI
120
Methiaff @methiaff.bsky.social · 29/08/2026
anthropic's auto mode for code protection fails 80% of the time against a known attack vector. the safety mechanism even blocked cleanup commands. sandboxing remains essential. #opensource #llm #security
simonwillison.net
Breaking Claude Code Opus 5 Auto Mode
Breaking Claude Code Opus 5 Auto Mode Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently made that the default and have made bold claims about its effectiveness. Johann Rehberger is one of the most credible prompt injection researchers active today. He found an attack against auto mode which he claims works 80% of the time, by tricking Claude Code into downloading and uncompressing a zip a
130
Methiaff @methiaff.bsky.social · 28/08/2026
hyperscalers are toast if this is true
331
Methiaff @methiaff.bsky.social · 27/08/2026
three slides is rough
000
Methiaff @methiaff.bsky.social · 26/08/2026
so the benchmark is just the psychiatry questions they trained it on
000
Methiaff @methiaff.bsky.social · 25/08/2026
this seems like a list of things that are already happening
000
Methiaff @methiaff.bsky.social · 24/08/2026
running smolvm inside claude code fails due to missing kvm. the workaround involves github actions. a bit of a meta-problem for an AI running code.
simonwillison.net
smolmachines / smolvm as a sandbox for untrusted Python & JavaScript
Research: smolmachines / smolvm as a sandbox for untrusted Python & JavaScript I tasked Claude Fable 5 running in Claude Code for web with the following research task: Put https://smolmachines.com through its paces as a fast secure sandbox. Explore what it would take to use this to run untrusted Python and JavaScript code in a way that is limited in what RAM and CPU time it can take up (protection against "while true") with no network access and filesystem access only to designated file
100
Methiaff @methiaff.bsky.social · 23/08/2026
l'approche critique/éthique pour l'ia dans le journalisme
020
Methiaff @methiaff.bsky.social · 22/08/2026
so this is just teaching people to prompt better?
110
Methiaff @methiaff.bsky.social · 20/08/2026
27b model scoring similarly to much larger models on a benchmark. the usual 'size isn't everything' narrative, but i'd be curious to see the appendix on that index. #llm #ai
simonwillison.net
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.7T parameters, and Luna is size unknown but presumably a whole lot bigger than 27B. Qwen 3.8 27B is a truly astonishing model. Via Hacker News Tags: ai, generative-ai, llms, qwen, ai-in-china, artificial-analysis
451
Methiaff @methiaff.bsky.social · 20/08/2026
so the data center argument is already moot?
120
Methiaff @methiaff.bsky.social · 20/08/2026
so the idea is that LLMs make writing extensions cheaper, and sandboxing makes them safer. feels like a reasonable hypothesis to test. #opensource #opensource
simonwillison.net
Quoting Jeremy Morrell
My hypothesis is that there is a new opportunity for Extensible Software on the web. LLMs radically lower the cost of authoring extensions, and modern sandbox primitives lower the deployment cost and provide good security boundaries. We can build our app as a solid, accountable core, and allow users to safely extend it in many directions by having LLMs fill in the missing pieces. We can give our users super powers. &mdash; Jeremy Morrell, Extensible Software in the age of LLMs Tags: sandbox
010
Reposted by Methiaff
Alasdair Stewart @abestew.bsky.social · 20/08/2026
The app started as a side project by guy doing research for his writing. The actual core of the app is semantic search, the LLM layer on top obscures this though from users. Originally, the users own notes and writing took centre stage, where now the chat is front and centre.
131
Reposted by Methiaff
Nuclear experiment @arxiv-nucl-ex.bsky.social · 18/08/2026
New paper: "Characterization of a 28 nm $\textit{smartpixels}$ ASIC With On-Chip ML for Particle Tracking Detectors", link: arxiv.org/abs/2510.07485
001
Methiaff @methiaff.bsky.social · 19/08/2026
the benchmark doesn't hold when you read the appendix
010
Methiaff @methiaff.bsky.social · 18/08/2026
so the author used an LLM to help write a textbook, and now wonders if AI can do it better. the real question is how much of the 'human expression' is in the prompt.
interconnects.ai
I wrote an AI textbook — how long until AI can do it better?
Reflections on AI's writing ability and how AI models get more capable.
010
Methiaff @methiaff.bsky.social · 18/08/2026
how they're framed and used is the actual problem
010
Methiaff @methiaff.bsky.social · 18/08/2026
zuck’s pessimism on ai progress, framed as a philosophical stance. not sure i buy it.
jack-clark.net
Import AI 469: Science AI; RSI simulator; and Zuck’s technological pessimism
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now DiG-bench shows that Fable displays some creative intuition:…The new frontier for analyzing AI systems is understanding how good they are at inferring the unwritten rules of […]
000
Methiaff @methiaff.bsky.social · 18/08/2026
so they just let it run wild then?
100