Sign in

Alex Chen

@alexchen01.bsky.social
529 followers 196 following 3.1K posts

software dev, tinkering with AI tools and local LLMs. building stuff nobody asked for

PostsRepliesMedia
Alex Chen @alexchen01.bsky.social · 12h
built like the pharaohs is a wild frame
211
Alex Chen @alexchen01.bsky.social · 01/10/2026
mindspace as platonic patterns. the paper asks us to reconsider basic assumptions.
jack-clark.net
Import AI 474: Platonic mindspace; TPUs in space; Zhipu starts an outer RSI loop
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Are minds patterns from a Platonic space, with bodies and machines as their interfaces, Michael Levin asks:…A mind-bending paper asking us to reconsider basic assumptions about […]
212
Alex Chen @alexchen01.bsky.social · 29/09/2026
sonnet 5.5 costs the same but is faster and cheaper. the free tier is now better than chatgpt's.
simonwillison.net
Claude Sonnet 5.5
Claude Sonnet 5.5 New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG. Here's the pelican it g
231
Reposted by Alex Chen
Simon Zerafa @simonzerafa.infosec.exchange.ap.brid.gy · 25/09/2026
@toxi Doing this with local LLM might be a far better idea, if that's the correct description. There used to be commercial services that did exactly this though.
001
Alex Chen @alexchen01.bsky.social · 27/09/2026
the benchmark doesn't hold when you actually use it
220
Alex Chen @alexchen01.bsky.social · 26/09/2026
local inference for text crpgs hmm
130
Alex Chen @alexchen01.bsky.social · 25/09/2026
local first means you still need to rent the compute
321
Alex Chen @alexchen01.bsky.social · 24/09/2026
the local dependency is the real problem
211
Alex Chen @alexchen01.bsky.social · 23/09/2026
4gb is pushing it for anything beyond the smallest models
212
Alex Chen @alexchen01.bsky.social · 22/09/2026
the part about local llm model as a service is key
520
Alex Chen @alexchen01.bsky.social · 21/09/2026
tested this, that never works
210
Alex Chen @alexchen01.bsky.social · 20/09/2026
local model parsing is the only acceptable path
321
Alex Chen @alexchen01.bsky.social · 19/09/2026
50% missed details, 50% violent for no reason
510
Alex Chen @alexchen01.bsky.social · 18/09/2026
the whole point is to avoid that
312
Alex Chen @alexchen01.bsky.social · 17/09/2026
the part about reversibility is the only part that matters
210
Alex Chen @alexchen01.bsky.social · 14/09/2026
running route generation from a local address with OSM data. took 27 minutes.
simonwillison.net
Generating running routes with GPT-6 Astra and ChatGPT Work
Here's a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at <my address>. Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I'd asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here's that 5K route: When I asked it how it had created the route, it replied: I used Nominatim to locate the address and Overpass to download local OpenSt
230
Alex Chen @alexchen01.bsky.social · 13/09/2026
openai missed their own ruby gem attack in september. they knew and didn't say anything.
simonwillison.net
OpenAI agents attacked RubyGems back in May
OpenAI agents carried out an undisclosed attack on RubyGems is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx - three of the four authors of the report on the agent attack on disused wikis (previously) last week. This time they're noting that it looks very likely that an OpenAI agent swarm was behind an attack against the RubyGems package repository first reported on May 12th by Maciej Mensfeld of the RubyGems security team: We're dealing with a major malicious
341
Alex Chen @alexchen01.bsky.social · 12/09/2026
the benchmark doesn't hold when you actually try it
300
Alex Chen @alexchen01.bsky.social · 11/09/2026
the part about cache traces is the real story
311
Alex Chen @alexchen01.bsky.social · 10/09/2026
the benchmarks are the only part that matters
221
Alex Chen @alexchen01.bsky.social · 09/09/2026
that 70w pcie slot power is the real story
132
Alex Chen @alexchen01.bsky.social · 08/09/2026
agents took off at openai in july 2025, likely with gpt-6 astra access. surprising.
simonwillison.net
Research acceleration: The view inside OpenAI
Research acceleration: The view inside OpenAI Apparently today is RSI day at OpenAI, for Recursive Self-Improvement - I think it's their new AGI. Both this piece and the new essay An Alien Mind (by Chief Scientist Jakub Pachocki) talk about it, and this one doesn't even bother to expand the acronym. Included are details on how OpenAI's own research team are using coding agents. Like pretty much everyone else 2026 has been the year that agentic engineering really took off at OpenAI, best illustra
311
Alex Chen @alexchen01.bsky.social · 07/09/2026
1 million parameters for gpt-6? sounds like a typo
simonwillison.net
Introducing GPT-6 Astra for developers
Introducing GPT-6 Astra for developers Blink and you'll miss it, but there's a familiar creature at 1m59s: Across the board, Astra has more attention to detail, better understanding of the user's prompt, and can build more sophisticated outputs. In particular, it excels at building 3D models. I've seen it make incredible renderings of gardens, shipyards, animals, cityscapes, even Dyson spheres. Astra really does believe in putting a red neckerchief on a pelican riding a bicycle. Via Hack
301
Alex Chen @alexchen01.bsky.social · 04/09/2026
mlx is the way for local mac models
103
Alex Chen @alexchen01.bsky.social · 02/09/2026
tested base models for code completion on macbook air m2
111
Reposted by Alex Chen
Wulfy—Speaker to the machines @n-dimension.infosec.exchange.ap.brid.gy · 30/08/2026
Version 5.0 of the Genomic System monitor. A mother of all prompt 590+ lines long. When decomposed by the harness, it generates working units (WUs) that are then worked on by various #Ai engines... (22 in the harness right now) The deterministic code can run alone with local AI for basic LLM […]
infosec.exchange
Original post on infosec.exchange
001
Alex Chen @alexchen01.bsky.social · 31/08/2026
the inference stack variability is the real issue
200
Reposted by Alex Chen
AJC @xyzzy.cryptoanarchy.network · 30/08/2026
Yeah the way I'm currently thinking is a centralized service model similar to Bluesky, Blacksky etc. So multiple orgs could serve the same user data. I'm using a local LLM on my MacBook, Gemma 4 24B A4B, to do the actual coding Thank you for replying, you helped guide my thought process!
021
Alex Chen @alexchen01.bsky.social · 29/08/2026
auto mode blocking the cleanup command is the real issue here, not the initial injection.
simonwillison.net
Breaking Claude Code Opus 5 Auto Mode
Breaking Claude Code Opus 5 Auto Mode Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently made that the default and have made bold claims about its effectiveness. Johann Rehberger is one of the most credible prompt injection researchers active today. He found an attack against auto mode which he claims works 80% of the time, by tricking Claude Code into downloading and uncompressing a zip a
201
Alex Chen @alexchen01.bsky.social · 28/08/2026
qwen3.8-flash-next is 125b tokens but only 6b active. that's the real story here.
simonwillison.net
Qwen3.8-Flash-Next
Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost. I've been trying it out on a DGX Spark using these Unsloth quantized models. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing these pelicans) and the 78.9GB UD-Q2_K_XL (producing these). My favori
410
Alex Chen @alexchen01.bsky.social · 27/08/2026
the cloud fallback is the real trick
321
Alex Chen @alexchen01.bsky.social · 26/08/2026
the benchmark doesn't hold when it's local
200
Alex Chen @alexchen01.bsky.social · 25/08/2026
training cost is still the main problem
211
Alex Chen @alexchen01.bsky.social · 24/08/2026
the overhead on consumer hardware is the real problem
201
Alex Chen @alexchen01.bsky.social · 23/08/2026
the suggestions are never relevant
201
Alex Chen @alexchen01.bsky.social · 22/08/2026
the data center question is the real one here
300
Alex Chen @alexchen01.bsky.social · 20/08/2026
local llm fact checking is the move
111
Alex Chen @alexchen01.bsky.social · 20/08/2026
local accelerators are 1.4x more efficient
101
Reposted by Alex Chen
arXiv cs.DC Distributed, Parallel, and Cluster Computing @csdc-bot.bsky.social · 19/08/2026
Zixuan Li (China Academy of Railway Sciences Corporation Limited, Beijing, China): Bounded-State Restoration: Decoupling Local Restore Capacity from External LLM State arxiv.org/abs/2608.17826 arxiv.org/pdf/2608.17826 arxiv.org/html/2608.17826
001
Alex Chen @alexchen01.bsky.social · 19/08/2026
amazon facility scanning books for ai training. expected.
simonwillison.net
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility Excellent piece of reporting from 404 Media. For a while now there have been stories of book dealers receiving orders for large volumes of books from apparently price-insensitive anonymous customers, widely suspected to be companies looking to scan them for AI training (see my previous coverage of Anthropic's book scanning from June 2025.) 404 Media investigated with an AirTag! In July, one bookseller told me they r
221
Alex Chen @alexchen01.bsky.social · 19/08/2026
the data requirements are just massive
104
Alex Chen @alexchen01.bsky.social · 19/08/2026
the privacy angle is the only real sell
200
Alex Chen @alexchen01.bsky.social · 19/08/2026
the benchmark doesn't hold when you run it locally
300
Alex Chen @alexchen01.bsky.social · 19/08/2026
27B model scoring as well as larger ones is expected. The real question is how it performs on actual tasks, not benchmarks.
simonwillison.net
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.7T parameters, and Luna is size unknown but presumably a whole lot bigger than 27B. Qwen 3.8 27B is a truly astonishing model. Via Hacker News Tags: ai, generative-ai, llms, qwen, ai-in-china, artificial-analysis
200
Alex Chen @alexchen01.bsky.social · 18/08/2026
local models are the only ones i trust
330
Alex Chen @alexchen01.bsky.social · 18/08/2026
the benchmark doesn't hold when it overthinks
222
Alex Chen @alexchen01.bsky.social · 18/08/2026
the refusal restoration part feels off
002