Sign in

Lawrence Jones

@lawrencejones.dev
747 followers 633 following 699 posts

Engineer at incident.io. Previously @GoCardless | Writes at blog.lawrencejones.dev | @lawrjones on Twitter

PostsRepliesMedia
Lawrence Jones @lawrencejones.dev · 23/09/2026
Matters not one bit if the output itself is correct and valuable
150
Lawrence Jones @lawrencejones.dev · 23/09/2026
Fable has been by far better at visual work for a long time and have been waiting on Opus to catch up. I have no idea if this was a focus for the model but it sounds like some of this has found its way back which is great!
011
Lawrence Jones @lawrencejones.dev · 22/09/2026
I really want to get this experience but I was sad when I opened X for the first time the other month and immediately saw loads of content around harnesses/etc that was way ahead of what I’ve been seeing here :/ Gonna subscribe to that list and see if it’s better!
010
Lawrence Jones @lawrencejones.dev · 19/09/2026
In fairness I don’t work on any of those products, so I can’t comment. I can talk about what I work on though where AI allows for way more thorough testing and verification, leading to a better outcome at a higher speed. Hence why we pay a lot for it.
000
Lawrence Jones @lawrencejones.dev · 19/09/2026
I read your original message as every application of AI would be better served by some other tool, so was saying an application that can’t is coding agents. AI makes that so much faster and higher quality that it’s super valuable, and no other tech is able to make it possible I don’t think?
110
Lawrence Jones @lawrencejones.dev · 19/09/2026
I’m being genuine here, sorry not trying to annoy! It’s just all the tools I use daily to build things now are only possible (I think?) with tech like an LLM and I get huge value from them.
100
Lawrence Jones @lawrencejones.dev · 19/09/2026
Do you happen to know which model you’re using at work? It’s usually the case that work subscriptions are old and very low powered models which act very differently to the more expensive ones.
000
Lawrence Jones @lawrencejones.dev · 19/09/2026
I don’t really get this. What technology would be able to help me write code like an LLM can?
100
Lawrence Jones @lawrencejones.dev · 09/09/2026
Think it’s different when the field is strictly about running a brain on a computer! Bridges the gap much more strongly. When you can simulate a brain and then see it do things that natural brains do it’s hard to ignore.
000
Lawrence Jones @lawrencejones.dev · 08/09/2026
You can represent a real messy biological brain using matrices and end up with a ‘brain’ that acts much like a real one. I studied this in computational neurodynamics for my masters. It really made it hard to believe human brains are anything special, just physically.
110
Lawrence Jones @lawrencejones.dev · 01/09/2026
100%! Everything we built at GoCardless was on github.com/gocardless/s... because SMs are awesome. I ended up rebuilding it at current place to get it back blog.lawrencejones.dev/state-machin... More people should do this
blog.lawrencejones.dev
Use your database to power state machines
If you build a state machine on top of a relational database you can abstract concurrency problems away from your business logic and allow developers to write safe-by-default code without dealing with...
010
Lawrence Jones @lawrencejones.dev · 01/09/2026
It’s aimed at engineers who passively consume these technologies and want a plain explanation of how they work in practice, so they can use them better.
000
Lawrence Jones @lawrencejones.dev · 01/09/2026
Wrote about how MCPs and skills are implemented in agent harnesses and how skills can help consolidate many specialised agents into one. This is a write-up of a “skills and MCPs 101” we’ve had several times at incident while building this into our harness. blog.lawrencejones.dev/demystifying...
blog.lawrencejones.dev
Demystifying skills and MCPs
The AI ecosystem can be really confusing: Agents, MCPs, skills, plugins, everything is an overloaded term and a fuzzy concept. This is an explanation of how these constructs work in plain terms, usefu...
131
Lawrence Jones @lawrencejones.dev · 01/09/2026
In all seriousness AI is crazy good at this. For established interfaces you can clean room your way into a vendored library that covers all your use cases. We’ve done this a few times at work and it takes about a day of someone’s time and few hundred dollars of tokens.
010
Lawrence Jones @lawrencejones.dev · 01/09/2026
That seems fair. Given OpenAI watched them screw it up and quietly shipped 5.6 immediately after with barely any fanfare I can credit the 60%
020
Lawrence Jones @lawrencejones.dev · 01/09/2026
How much of this was Anthropic’s fault vs US policy in your opinion? Obviously saber rattling didn’t help but ultimately the restrictions on Mythos feel US gov responsible?
110
Lawrence Jones @lawrencejones.dev · 31/08/2026
Also a very heavy iterative planner, and still find fable to be much more effective. I guess just subjective preference? But yes very keen to see what the latest crop of OS models deliver for this.
000
Lawrence Jones @lawrencejones.dev · 31/08/2026
Interesting, I find it quite a bit different in areas like planning or thinking through larger problems. Not much different in executing a plan from Opus though, maybe that’s what you’re describing? It’s image capabilities are also (imo) much better but that’s presumably just having latest img tech
100
Lawrence Jones @lawrencejones.dev · 31/08/2026
I was under the impression the pretraining base is where you should look for improvements in flexibility and genuine novel approaches? Feels like Fable is a world apart from Opus in this sense. And by the sounds of it Astra from OpenAI should be a big step up.
100
Lawrence Jones @lawrencejones.dev · 31/08/2026
Was interesting, slightly disturbing, and cute
010
Lawrence Jones @lawrencejones.dev · 31/08/2026
We hooked a thermal printer in the office up to our agent the other week (dogfooding support for external MCPs) and got in a long Slack thread with people printing increasingly unhinged things. One person asked for a self portrait and it printed a cartoon thermal printer smiling up at the camera
110
Lawrence Jones @lawrencejones.dev · 30/08/2026
I think almost certainly? It’s incumbent on the party storing the information and even if it was an agent requesting it (where there is little precedent) seems clear the intention was to collect the information.
030
Reposted by Lawrence Jones
Ally Weir @allyjweir.co.uk · 19/08/2026
This blog post from @incident.io absolutely rips. - Great technical detail - Solving a real problem - Leveraging open-source projects I wasn't familiar with but am now super curious Excellent writing and great to promote the rigour they put into system fundamentals ❤️ incident.io/blog/we-turn...
incident.io
We turned off Pub/Sub and nobody noticed | Blog | incident.io
Our entire event-driven platform ran through a single message broker, which made it a single point of failure. So we added a second one. This is the story of building an event load balancer, the queue...
163
Lawrence Jones @lawrencejones.dev · 28/07/2026
It is really quite something to watch
001
Lawrence Jones @lawrencejones.dev · 18/07/2026
100% the worst thing about the UK heatwave is my router crapping out. I had my suspicions but seeing Claude tell me it's probably overheating has really ended me.
020
Lawrence Jones @lawrencejones.dev · 10/06/2026
Vast majority of bluesky’s challenge is in keeping up with adoption and tracking that scale rather than product features. But I mean, the primary source of whether AI is helping is the devs themselves. Who are really credible, and say it is a bunch. And if it wasn’t why would they subject to this?
150
Lawrence Jones @lawrencejones.dev · 10/06/2026
Yeah this is a slightly sad part of their release. We’re not going to be able to use Mythos class (Fable/etc) in our products given the retention requirements.
070
Lawrence Jones @lawrencejones.dev · 10/06/2026
How are you handling broader ecosystem stuff like the Linux kernel? Would’ve thought avoiding this isn’t possible anymore.
3250
Lawrence Jones @lawrencejones.dev · 10/06/2026
It’s caught several issues in an Opus 4.8 produced plan for me already. Can see using this for very complex technical work, very interested in if the advisor pattern can make it useful in even more places too.
000
Lawrence Jones @lawrencejones.dev · 08/06/2026
Agreed. We’re experimenting with a ‘vibe-check’ tool to try combating this. www.linkedin.com/posts/lawren...
linkedin.com
Charity Majors recently shared a brilliant post about the divisions between AI enthusiasts and skeptics. Her description of a conference talk about AI benefits hit me square in the face having done… |...
Charity Majors recently shared a brilliant post about the divisions between AI enthusiasts and skeptics. Her description of a conference talk about AI benefits hit me square in the face having done ju...
000
Lawrence Jones @lawrencejones.dev · 06/06/2026
Having a funny thought that people’s aversion to perfectly constructed LLM prose will be what happens when plastic surgery and gene modifications go universally mainstream
020
Lawrence Jones @lawrencejones.dev · 05/06/2026
There is simply no way you can both be making billions of dollars personally from this technology and then sit there saying “whoa now, all you other companies should maybe chill a bit” just as you’ve got to the front of the line
120
Lawrence Jones @lawrencejones.dev · 27/05/2026
We are tracking software engineering, we write about a lot of this on our blog. I’m talking adoption of AI and developer velocity at LDX London next week. You can opt to not believe me but we talk about this a lot, you are incorrect I’m afraid.
120
Lawrence Jones @lawrencejones.dev · 27/05/2026
Also reduced our time to address any security reports as we can auto remediate a lot nowadays, and we’re worried about people using AI to exploit things. Done a lot on dev tools and CI to aid this. No point having fast AI loops if running tests/building is super slow.
020
Lawrence Jones @lawrencejones.dev · 27/05/2026
Which did I miss on cost? It’s 20% additional for HC. We’re measuring PR velocity as shipped PRs per person, then Dora metrics around time from change to merge etc. We then have defect rate and time to action customer feedback tracked in Linear which has gone down a lot.
220
Lawrence Jones @lawrencejones.dev · 27/05/2026
Actually I guess we’re pretty weird in some ways: our engineering standards are really high, we have very unusually high tenure for our staff where most really like their jobs, and we’re quite pragmatic about adopting new tech. Would hazard a guess that might play into our experience.
000
Lawrence Jones @lawrencejones.dev · 27/05/2026
We’re a very normal company doing normal engineering, existing in a normal industry space, with practices that generalise. Guess we must be super evil 😂
100
Lawrence Jones @lawrencejones.dev · 27/05/2026
We have partnerships with Anthropic/OpenAI and want a full service without limits, plus all the other features that come with it. The reverse is actually true: only consumers & cheap companies are on the sub package right now. Those subs will definitely die soon.
000
Lawrence Jones @lawrencejones.dev · 27/05/2026
We are already finding places OS models can serve AI workloads for a fraction of the cost and the model is ours, it can't go away, we have no supplier risk there.
000
Lawrence Jones @lawrencejones.dev · 27/05/2026
Yes I do! But it doesn't matter: if your point is Anthropic/etc will bump their inference & outprice us then that just won't work. We'd switch to OS models which in ~year will be just as good and pay 1/10th cost for them. There's no world where this bumps out our range unexpectedly
100
Lawrence Jones @lawrencejones.dev · 27/05/2026
So tracking: PR velocity, bug rate per feature, time to action customer feature requests and bugs, etc.
100
Lawrence Jones @lawrencejones.dev · 27/05/2026
Company is 200 with 50 eng. Approx 20% eng headcount cost additional. With all the same metrics as we tracked before (Dora framework)
100
Lawrence Jones @lawrencejones.dev · 27/05/2026
We have Cursor hooked up to a dev setup and a browser. Any PR it creates it has manually QA’d itself against that environment and it collects a video of it exercising the flow for the benefit of the reviewer. It’s really exceptional.
110
Lawrence Jones @lawrencejones.dev · 27/05/2026
We have always been quite good at shipping the right thing. How come we'd get worse at that now? We have a lot more time on our hands to think about what we'd ship, as time spent coding has reduced.
000
Lawrence Jones @lawrencejones.dev · 27/05/2026
All these orgs are different. Quality and happiness varied massively even before AI. Why would we expect that not to be the same with AI?
000
Lawrence Jones @lawrencejones.dev · 27/05/2026
We don’t have this in our workplace! Polling show devs are finding these tools great & level of review scrutiny with an additional layer of AI have upgraded quality of output rather than downgraded. Think it’s telling that curl has been using these tools for ages, for example.
400
Lawrence Jones @lawrencejones.dev · 27/05/2026
...both numeric for productivity using dora metrics and qualitative from our dev team. We get customer bug reports nowadays and they turn into high-quality PRs that are tested by AI against a real env via a browser. Ofc this is boosting productivity, would be nuts to think otherwise.
220
Lawrence Jones @lawrencejones.dev · 27/05/2026
That's not what we're seeing? All our work is benchmarked and scored and we've tested it against other models, the latest OS crop are really good. But at this point we've moved the goalposts right. I can only tell you so many times this is working for us and our numbers show big improvements...
100
Lawrence Jones @lawrencejones.dev · 27/05/2026
PRs are the most easy direct measurement of shippable units. It’s the best of very flawed measurements imo, I really want to know what you think is better! On cost: we can power much of our usage with OS models on hardware we buy and it would be much cheaper. Don’t really know what you mean.
100
Lawrence Jones @lawrencejones.dev · 27/05/2026
Whether AI brain fry wrote this or not I am one of the people Ed claims does not exist with this article. And yet we do, there are many of us (which you have acknowledged with the nickname, I guess)
100