Sign in

Chris Parsons

@chrismdp.com
477 followers 834 following 1.7K posts

I help leaders leverage AI in their organisations - see chrismdp.com 💪 also co-founder & CTO of cherrypick.co

PostsRepliesMedia
Chris Parsons @chrismdp.com · 16h
The most important thing about any system is figuring out how to record structured feedback from the people who understand it, so that an AI can continually improve the system. www.chrismdp.com/prompt-evals-are-u…
chrismdp.com
Prompt Evals Alone Are Useless
Over the last couple of nights I have been building a text adventure RPG, mostly to see what Astra could do. I have been giving the AI an end of day prompt, ...
000
Chris Parsons @chrismdp.com · 20h
And it's better. www.chrismdp.com/time-to-hire-an-ai…
chrismdp.com
Time To Hire An AI Head
For a long time I have wanted to hand AI a senior job. When I was a CTO, this was the kind of work I would have delegated to a head of department: take a loo...
000
Chris Parsons @chrismdp.com · 20h
I normally run about 3-5 AI sessions at once with myriad subagents and several constant streams of work, often through the evening and into the small hours for me to review in the morning. Astra 6 burned through a Max 20 after 18 hours; Opus 5.5 has been running for 3 days and has used only 57%.
100
Chris Parsons @chrismdp.com · 20h
I cannot believe how much usage I'm getting from Opus 5.5 compared to Astra.
100
Chris Parsons @chrismdp.com · 29/09/2026
This seems good. I imagine they were desperate to get it out, with Opus 5.5 already out. Holding it back suggests they’re taking safety seriously. krdo.com/money/cnn-business-consume…
krdo.com
‘Didn’t quite meet the bar’: OpenAI won’t release new AI model due to safety concerns
By Auzinea Bacon, CNN (CNN) — OpenAI said it won’t release its latest model, dubbed GPT-6.1 Astra, because it “didn’t quite meet the bar” for
000
Chris Parsons @chrismdp.com · 29/09/2026
OpenAI has shelved GPT-6.1 Astra. Its safety lead says the model fell short on staying within its authorised scope and telling users what it had done.
100
Chris Parsons @chrismdp.com · 29/09/2026
I now do 95% of my work in Claude, WhatsApp and Telegram. www.chrismdp.com/claude-code-on-you…
chrismdp.com
Run Claude Code From Your Phone
Your laptop is a terrible development environment. It sleeps when you close the lid, loses state when you reboot, and chains you to a single device. The solu...
001
Chris Parsons @chrismdp.com · 29/09/2026
Karpathy playlist: youtube.com/playlist?list=PLAqhIrjk… 3blue1brown series: www.youtube.com/watch?v=aircAruvnKk…
youtube.com
Neural Networks: Zero to Hero
Andrej Karpathy on YouTube
000
Chris Parsons @chrismdp.com · 29/09/2026
Anyone with a software engineering background and roughly A-level / high school maths can follow it, no data/ML experience required. 3blue1brown's YouTube series on LLMs is also amazing, more of a detailed explanation rather than a walkthrough.
100
Chris Parsons @chrismdp.com · 29/09/2026
Want to learn exactly how LLM work? I'd highly recommend Karpathy's videos building an LLM from scratch.
111
Chris Parsons @chrismdp.com · 28/09/2026
Which of your AI models are you still managing yourself? www.chrismdp.com/time-to-hire-an-ai…
chrismdp.com
Time To Hire An AI Head
For a long time I have wanted to hand AI a senior job. When I was a CTO, this was the kind of work I would have delegated to a head of department: take a loo...
000
Chris Parsons @chrismdp.com · 28/09/2026
What I don't do any more is manage the juniors myself. When my plans ran out last week I tried it, and was straight back to feeding context and pressing approve all day. Six patterns in the post, from sending the Head to visit the expert, to asking it to explain strategy while you keep the call.
100
Chris Parsons @chrismdp.com · 28/09/2026
The juniors are open models like GLM 5.3 Flash. They're finally competent and cost pennies, but they still miss stuff. So the Head insists on tests that check the real requirement, not the junior's word that it's done.
100
Chris Parsons @chrismdp.com · 28/09/2026
Here's how I'm doing it safely: I hand the whole outcome to a Head: a strong model like Astra, or Opus 5.5 when I want to save money. It fetches its own context, splits the job up, briefs cheaper models and owns whether the work is accepted.
100
Chris Parsons @chrismdp.com · 28/09/2026
I have found the latest AIs are good enough for senior work.
100
Chris Parsons @chrismdp.com · 27/09/2026
indeed - we still haven't figured out the right tools here either. Cowork/Codex just aren't there yet
110
Chris Parsons @chrismdp.com · 27/09/2026
Luckily there are now OpenRouter endpoints, and a whole host of new decision models coming down the tracks, so I'm able to work on putting these narrow use cases in to apps. www.chrismdp.com/ai-decisions-just-…
chrismdp.com
Redirecting…
000
Chris Parsons @chrismdp.com · 27/09/2026
Is Jev worth the hype? Yes, actually. It's going to be extremely useful for glueing AI systems together in a much more efficient and cheaper way, and it means we can take on a new class of problems.
100
Chris Parsons @chrismdp.com · 27/09/2026
For example, if there’s a classification step you need a “other” or “none of the above” as an escape hatch when the data it is examining is irrelevant or not applicable. If you don’t do this you’ll get nonsense.
210
Chris Parsons @chrismdp.com · 27/09/2026
But on anything that needs open judgement it was no better than the cheapest LLM, and sometimes worse, depending on how you phrase the question.
100
Chris Parsons @chrismdp.com · 27/09/2026
On closed-set classification (routing, triage, category assignment) Jev matched or beat the language models I already use, far faster and far cheaper.
100
Chris Parsons @chrismdp.com · 27/09/2026
Some intel on where Jev might fail to work for you (and why that doesn't always matter):
110
Chris Parsons @chrismdp.com · 26/09/2026
Productive individuals do not make productive firms. AI rollout work has to become shared operating practice, or it stays as private cleverness.
131
Chris Parsons @chrismdp.com · 25/09/2026
I taught a bunch of designers how to think like CTOs for getting their apps going yesterday. www.chrismdp.com/ship-like-an-engin…
chrismdp.com
Ship Like an Engineer Without Writing Any Code
This post is based on a workshop given on 24th September 2026 at design+AI.Building a working prototype with AI is easy now: ask Claude to build you somethin...
000
Chris Parsons @chrismdp.com · 25/09/2026
We still need engineers more than ever to supervise the hard stuff, but the way to make that point is not to prevent others from getting into code with AI tools: with a few principles anyone can use AI to ship scalable, tested apps.
100
Chris Parsons @chrismdp.com · 25/09/2026
I can't stand the coder gatekeeping I see everywhere! Engineering is no longer out of reach with AI.
110
Chris Parsons @chrismdp.com · 24/09/2026
all set up for my Design+AI workshop today! Teaching designers how to think like an engineer in their production projects - I’ll post a proper write up of session soon
031
Chris Parsons @chrismdp.com · 23/09/2026
www.chrismdp.com/the-harness-is-the…
chrismdp.com
The Harness Is The Bottleneck
DeepSeek V4-Flash 0731 is doing better than I expected.It does not beat the best premium coding models: it is slower, stops early more often, and needs a dec...
010
Chris Parsons @chrismdp.com · 23/09/2026
I have a myriad of other agents, checking prospect + client status, managing my content queue, and matching receipts + invoices. I almost never check email any longer. I don't worry about my business or personal accounts - reconciling takes about five minutes.
100
Chris Parsons @chrismdp.com · 23/09/2026
My email agent Em noticed that Opus 5.5 and Sol/Terra 6 were out via a newsletter. She forwarded the info to Richard my research agent, who compared them, looked up references in my file-based wiki and saved references.
100
Chris Parsons @chrismdp.com · 23/09/2026
Earlier this month I left the country for four days. My AI kept the business running while I was gone.
100
Chris Parsons @chrismdp.com · 22/09/2026
Can't take credit for first coining this one, but hope to see it become more popular! www.chrismdp.com/call-it-machine-in…
chrismdp.com
Call It Machine Intelligence
Donald Trump wants a better name for artificial intelligence. On Saturday he called the words inaccurate and very ineloquent, and asked his followers to vote...
000
Chris Parsons @chrismdp.com · 22/09/2026
I suggested "Machine Intelligence" name in a WhatsApp group with other AI experts, and I haven't heard a better one. It covers any intelligence produced by machines, and says nothing about real or fake or about biology.
100
Chris Parsons @chrismdp.com · 22/09/2026
"Artificial" implies that the intelligence isn't real. Biology has had the monopoly on intelligence until now, so we've conflated the two for our entire existence. That monopoly is over.
100
Chris Parsons @chrismdp.com · 22/09/2026
John McCarthy, who coined the term "Artificial Intelligence", later said he wished he'd called the field "computational intelligence". That's no better than "artificial": it's long and awkward to say.
100
Chris Parsons @chrismdp.com · 22/09/2026
He asked his followers to vote between Superior, Extreme and Supreme Intelligence (and then dropped Supreme).
100
Chris Parsons @chrismdp.com · 22/09/2026
Trump is right: "artificial intelligence" is the wrong name. "Machine Intelligence" is much better than his options though.
100
Chris Parsons @chrismdp.com · 21/09/2026
It looks like a bet on a) ever increasing AI capability is possible given more compute and b) demand keeps up with the capability on offer within the investment cycle. Remains to be seen whether this is a good bet: seems high risk to me
100
Chris Parsons @chrismdp.com · 20/09/2026
How about an info-sharing deal with China to allow them to catch up to Astra levels in return for slowing further work, allowing UN inspectors to AI labs worldwide, and share a Nobel peace prize? www.chrismdp.com/truman-would-have-…
chrismdp.com
Truman Would Have Taken Over AI by Now
Anthropic announced the most powerful cyberweapon ever built last week and kept it, granting access to forty companies while the US government got a press re...
000
Chris Parsons @chrismdp.com · 20/09/2026
Renaming Artificial Intelligence is definitely not the most important thing on the US president's agenda.
100
Chris Parsons @chrismdp.com · 20/09/2026
www.chrismdp.com/build-an-ai-knowle…
chrismdp.com
Build An AI Knowledge Base From Scratch
What if the AI already knew what you were working on? What if it remembered what you decided last week, what your board presentation covers, what your team s...
000
Chris Parsons @chrismdp.com · 20/09/2026
I've had great success with adding an automatic transcript check before adding the transcript to your personal file vault: given the context of the vault, it will often correct misspellings and prevent the proliferation of bad information.
100
Chris Parsons @chrismdp.com · 20/09/2026
www.chrismdp.com/prompt-evals-are-u…
chrismdp.com
Prompt Evals Alone Are Useless
Over the last couple of nights I have been building a text adventure RPG, mostly to see what Astra could do. I have been giving the AI an end of day prompt, ...
000
Chris Parsons @chrismdp.com · 20/09/2026
Karpathy's autoresearch and Shopify's pi-autoresearch climb a single number. A story has no such number. Whether it is good is a judgement, so the judge has to be an agent reading the whole run. If your AI team only shows you prompt scores, ask what tests the code, memory and screen.
210
Chris Parsons @chrismdp.com · 20/09/2026
Think of the top layer as BDD end to end scenarios from my Cucumber days, applied to something far more subjective.
100
Chris Parsons @chrismdp.com · 20/09/2026
It started with the model and stopped there. It now has three layers: ordinary software tests for rules and persistence, replay cases for one saved call, and whole-system scenarios that exercise code, memory and the screen.
100
Chris Parsons @chrismdp.com · 20/09/2026
My AI eval stack was upside down.
100
Chris Parsons @chrismdp.com · 19/09/2026
Really impressive demo from Ethan Mollick of using Blender to create a storyboard and then upload that to a video generator - reminds me of my using Claude Design but on steroids www.chrismdp.com/software-factories…
000
Chris Parsons @chrismdp.com · 19/09/2026
www.chrismdp.com/ai-progress-is-not…
chrismdp.com
AI Progress Is Accelerating. Here Is Why It Feels Slow
AI progress still compounds through longer task horizons, better tools and lower costs. See which measures leaders should track beyond benchmark hype.
010
Chris Parsons @chrismdp.com · 19/09/2026
I also kept the honest part on the page: 50% success is too weak for most unattended work, and METR says estimates above 16 hours are unreliable. The three measures that matter for leaders are task horizon, cost per accepted completion, and human review required.
100