Sign in

Jim Bennett

@jimbobbennett.dev
1.2K followers 842 following 972 posts

World's most energetic dev rel Microsoft MVP. 🌈ally. I ❤️ Star Wars Lego & 🐻‍❄️. Father, husband. He/him.

PostsRepliesMedia
Jim Bennett @jimbobbennett.dev · 04/09/2026
I keep hearing folks say "delete all your skills" when a new model comes out. No - skills can be pipelines of work, not just "this is how you interact with a service". Audit, eval, don't just blindly delete. I wrote about this here: jimbobbennett.dev/blogs/skills...
jimbobbennett.dev
Don't delete your skills, audit them
Boris Cherny says to delete your CLAUDE.md, skills, and hooks every six months. That's good advice for one kind of skill and dangerous for another. Here's how to tell them apart, and why the skills th...
132
Jim Bennett @jimbobbennett.dev · 26/08/2026
Oh yeah! This is a LOT of fun!
021
Jim Bennett @jimbobbennett.dev · 07/08/2026
How do you trust agents before they hit production? At @arize.bsky.social's Observe 2026, Salesforce's Manjit Singh walked through evaluating agents across the lifecycle — single agents to orchestrators & handoffs. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 05/08/2026
"Everyone is an AI agent builder — if you let them." At @arize.bsky.social's Observe 2026, CrewAI shared enterprise lessons on where agent ROI shows up and how to scale building across an org. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 03/08/2026
As agents become digital coworkers, access control built for humans starts to break. At @arize.bsky.social's Observe 2026, WorkOS's Michael Grinich explored the identity & security model for autonomous agents. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 31/07/2026
What does agent adoption actually look like in production? At @arize.bsky.social's Observe 2026, Mastra shared patterns from thousands of teams — what separates shipped agents from stuck prototypes. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 29/07/2026
From experimentation to production, agents need whole-lifecycle platforms. At @arize.bsky.social's Observe 2026, Microsoft's Sebastian demoed building, deploying, evaluating & governing agents with Microsoft Foundry. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 27/07/2026
Scaling agents from prototype to production is an infrastructure problem. At @arize.bsky.social's Observe 2026, Anyscale's Robert Nishihara explained how Ray scales RL, inference & multimodal AI — and why RL is having a moment. Sketchnoted 👇 🔗 Video link in the comments.
220
Jim Bennett @jimbobbennett.dev · 24/07/2026
The hardest problems in AI aren't model problems anymore — they're evaluation problems. At @arize.bsky.social's Observe 2026, Hamel Husain argued agents are bringing the data scientist back to AI engineering. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 22/07/2026
What are the smartest AI investors seeing before everyone else? At @arize.bsky.social's Observe 2026, Jaya Gupta of Foundation Capital shared where VC is flowing across the AI stack. Sketchnoted the fireside 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 20/07/2026
When a hallucination is a regulatory + financial risk, responsible AI gets real. At @arize.bsky.social's Observe 2026, BlackRock shared how it deploys AI to support pros managing trillions — with real guardrails & evaluation. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 17/07/2026
With autonomous agents, observability shifts from "what happened" to "why did the agent do that." At @arize.bsky.social's Observe 2026, AWS's Nate Slater explored how agentic AI changes observability. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 15/07/2026
Two of the fastest-growing open-source agent projects, one conversation. At @arize.bsky.social's Observe 2026, OpenClaw & Nous Research debated where agent frameworks go next — memory, skill creation, long-term learning. Sketchnoted 👇 🔗 Video link in the comments.
120
Jim Bennett @jimbobbennett.dev · 13/07/2026
"AI agents need specs, not prompts." At @arize.bsky.social's Observe 2026, George Zhang argued the engineer's real job is specifying the hill agents climb — tests, evals, rubrics, constraints. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 10/07/2026
What happens when a leading AI coding company turns its product inward? At @arize.bsky.social's Observe 2026, Cursor shared how it uses agents, evals & agent-powered workflows to build Cursor itself. Sketchnoted 👇 🔗 Video link in the comments.
130
Jim Bennett @jimbobbennett.dev · 08/07/2026
"Kubernetes is not your sandbox." At @arize.bsky.social's Observe 2026, the Daytona team argued K8s wasn't built for agent workloads, and walked through what agent-native infrastructure actually needs. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 06/07/2026
Building agents is harder than the demos make it look. At @arize.bsky.social's Observe 2026, Anthropic's Marius Buleandra shared why agent failures compound in production, how to design evals that catch them, and why human review still matters. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 03/07/2026
How do you improve a product with hundreds of millions of users? At @arize.bsky.social's Observe 2026, OpenAI's Stuart Sy showed how ChatGPT turns fragmented 'vibes' into evidence + action. I sketchnoted the talk 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 01/07/2026
Agents stopped being demos this year — they're shipping code, fixing bugs, and running real workflows. In @arize.bsky.social's Observe 2026 keynote, the founders lay out what's next. I sketchnoted the whole keynote 👇 🔗 Video link in the comments.
100
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 24/06/2026
A field guide to four of the latest from @jimbobbennett.dev, where each one leaks, and why no single pass rate was ever going to survive this. arize.com/blog/long-h...
011
Jim Bennett @jimbobbennett.dev · 15/06/2026
Do you have an AI agent? Do you actually know what it is doing? Do you know if it works? Typically the answer to the first question is yes, and for the second it's we think so, based off 'vibes'. Which is a terrible way to build and run production software. 1/2
110
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 12/06/2026
Apple paid Google ~$1B/yr to license memory for Siri. OpenAI rebuilt ChatGPT memory in place. Anthropic gave models an API to consolidate their own. All called "memory." None is what users mean. @jimbobbennett.dev wrote a field map: arize.com/blog/memory...
arize.com
Memory is still a missing primitive: Cataloguing what the field is actually shipping
This week the field shipped four kinds of memory, and Apple paid Google a billion dollars a year for one of them. None of the four is what the demos imply. A field map of what's actually shipping, and the missing primitive that sits between the buckets.
012
Jim Bennett @jimbobbennett.dev · 08/06/2026
"I genuinely don't care. Pick one." That was my contribution to a meeting last week where the team was debating two tools. And it was the most useful thing I said all day. "Strong opinions, loosely held" is the "approved" take. I think it's mostly nonsense.
linkedin.com
Strong opinions, strongly held - and why I don't care about your tooling debate
I was in a meeting last week where the team was debating which of two tools to use for a job. Both of them do the thing.
110
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 03/06/2026
Microsoft picked OpenInference. Twice. The open trust stack for AI agents announced at #MSBuild, ASSERT for evaluation, ACS for controls, both ride on the open tracing standard Arize built for agents. arize.com/blog/micros...
arize.com
Microsoft's open trust stack runs on OpenInference
Microsoft's open trust stack for AI agents puts ASSERT and Agent Control Specification on top of OpenInference, connecting evaluation, runtime controls, and observability through a shared trace contract.
011
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 02/06/2026
At Microsoft Build? Our 2 must do things for today: 1. Catch Sarah Bird's session - Observe and control agents with OSS tools build.microsoft.com/en-US/sessi... 2. Head to the Microsoft AI expert booth to meet with @jimbobbennett.dev from our devrel team about AI observability and Evals #MSBuild
build.microsoft.com
Observe and control agents across any framework with open source tools
As AI agents move into production, developers own safety, governance, and reliability across Microsoft Agent Framework and open-source stacks. This session shows how to govern agents end to end: turning your requirements into context-aware evaluations, stress-testing against adversarial risks, applying open controls that work across frameworks, and keeping humans in the loop on high-stakes actions. Leave with a blueprint for shipping agents at enterprise scale. Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
011
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 01/06/2026
Will you be at Microsoft Build this week, either in person in SF or virtually? Our very own @jimbobbennett.dev will be giving a demo session on understanding and fixing agents with open source observability and evals, Wednesday 3:30pm, Theater C. #MSBuild build.microsoft.com/en-US/sessi...
build.microsoft.com
Understand and fix Agent Framework apps with observability and evals
Your AI apps are getting more complex, with multiple agents, tools, and different orchestration patterns. This makes them harder to understand, debug, and test. This session shows you how to visualize the decisions your LLMs are making in complex Microsoft Agent Framework applications, using open standards and open source tooling to provide you with instrumentation and observability. You'll also see how you can use another LLM as a judge to evaluate how well your AI app is working. Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
011
Reposted by Jim Bennett
arize-phoenix @arize-phoenix.bsky.social · 21/05/2026
Phoenix now lets you compose evaluation strategies in code. Most eval tooling hands you a fixed menu of judge templates. Real evaluation is rarely that tidy.
131
Jim Bennett @jimbobbennett.dev · 20/05/2026
All Londoners who read this post will be reminded of it every day when they get on the tube…
000
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 20/05/2026
Your AI agent disagrees with your human reviewers all day. Most teams treat that as noise. It's the most useful signal in the system. @jimbobbennett.dev wrote up how to mine the gap and feed it back to the agent. arize.com/blog/self-i...
011
Jim Bennett @jimbobbennett.dev · 19/05/2026
Every AI agent deployed inside an enterprise sometimes quietly disagrees with the humans running the same process. The written policy says one thing. The institutional knowledge sitting in Slack threads, hallway conversations, and the heads of long-tenure employees says another. 🧵 1/3
100
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 12/05/2026
One AI Question with Cam Young We asked our Strategic AI Solutions Architect: What's a 🔥 take on evals? His answer: Stop guessing and start measuring. Use "LLM-as-a-judge" for nuance, but don't ignore code-based evals for speed and human annotators for ground truth. #AI #AIStrategy #AIEvals #LLM
011
Jim Bennett @jimbobbennett.dev · 12/05/2026
Claude Code users - want to be notified when Claude wants your attention? If you have a RPi and a 3.5" screen, then here's a project that puts a happy character on the screen. Bored when Claude is busy, dances when Claude needs your attention. All the code is here: github.com/jimbobbennet...
github.com
GitHub - jimbobbennett/claude-notify: Raspberry Pi + 3.5" touchscreen Claude mascot that dances when Claude Code on your Mac needs your input
Raspberry Pi + 3.5" touchscreen Claude mascot that dances when Claude Code on your Mac needs your input - jimbobbennett/claude-notify
131
Reposted by Jim Bennett
Laurie Voss @seldo.com · 08/05/2026
How many instructions can you give an LLM before it starts to forget about some of them? How long can your skill file be? How big can your prompt get? I did some actual *research*! www.linkedin.com/pulse/models-got-o…
Graph showing accuracy vs constraint keywords (N) for 6 AI frontier models. As constraints increase from 10 to 10,000, accuracy generally decreases. GPT-5.5 and Gemini 3.1 Pro maintain highest performance, while older models (GPT-4.1, Claude Sonnet 4) decline faster.
2406
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 07/05/2026
🚀 One AI Question with Aparna Dhinakaran We asked our Chief Product Officer: When should I start doing evals? Her answer: Start now. Don't wait—look at your data and traces immediately to find where your agents fail. #AIOps #LLMOps #AIEvals #MachineLearning #AI
011
Jim Bennett @jimbobbennett.dev · 07/05/2026
3 strikes and you’re an AI skill. My thoughts on what AI skills are and converting regularly used prompts to skills. jimbobbennett.dev/blogs/3-stri...
jimbobbennett.dev
3 strikes and you're an AI skill
Back in the day when we wrote actual code instead of poking at an AI, I had a general rule for when to refactor repeated code. Do it once, fine. Do it a second time, fine. Do it a third time - refacto...
010
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 30/04/2026
One AI Question with @jimbobbennett.dev What's your 🌶️ take on AI? Our DevEx Engineer's take: Start with the mindset that AI sucks—so you're forced to build the evals and observability to make it great. Don't trust it. Test it. #AI #Programming #SoftwareDevelopment
121
Jim Bennett @jimbobbennett.dev · 24/04/2026
Why do I run? Currently I’m running for Sage House, a dementia support charity helping my dad and other folks live with dementia. Sponsor me for the @londonmarathon.bsky.social! 2026tcslondonmarathon.enthuse.com/pf/james-ben...
000
Reposted by Jim Bennett
pamelafox.bsky.social @pamelafox.bsky.social · 14/04/2026
I'm excited that @seldo.com from Arize will be speaking at AgentCon about one of my fav topics, evals: "Stop vibe-testing: run real agent evals" Join us on May 4th at the Computer History Museum in Mountain View: luma.com/u96hax55
luma.com
AgentCon - Silicon Valley · Luma
Join us for the AI Agents World Tour, a global series of one-day conferences designed exclusively for developers building the future with AI agents. Confirmed…
012
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 06/04/2026
Last week at the AI Builders meetup in Seattle, @jimbobbennett.dev spoke about boosting Claude Code performance using Prompt Learning. If you missed it, don't worry - Jim recorded a video of the session. youtu.be/ES43SEXArvk
youtube.com
Boost Claude Code performance with prompt learning - optimize your prompts automatically with evals
🚀 Boost the performance of Claude Code with prompt learning🧠 Repo with all the code here: https://github.com/Arize-ai/prompt-learningPrompt Learning, a new...
001
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 19/03/2026
Part 2 of our deep dive into how we built Alyx: context windows arize.com/blog/how-to... Once an agent starts running, context becomes the bottleneck fast.
arize.com
How We Keep Alyx's Context Window From Steamrolling Itself
A deep dive into LLM context management: middle truncation, external memory, deduplication, and sub-agent architectures.
121
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 16/03/2026
TDD doesn't work for AI apps, instead you need EDD - eval-driven development. @jimbobbennett.dev gave a talk on this very topic at Azure AI Connect, part of the Global AI Community. Check out the video to learn how to do EDD, and for some very 🌶️ takes on AI. www.youtube.com/live/MC8yKi...
youtube.com
Azure AI Connect 2026 Day 5
The Future of AI is Connected.The Future is on Azure.Join us for a free 5-day virtual event dedicated to mastering the Microsoft Azure AI platform.March 2-6,...
022
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 16/03/2026
Add instrumentation to your #AI apps in 1 terminal command and 1 prompt! @jimbobbennett.dev put together this video to show you how, using our newly released skills for your favorite coding agent. youtu.be/qby0FKv-IfA
youtube.com
Arize Skills: Add Instrumentation & Tracing to Your AI App with Claude Code, Copilot, or Cursor
🚀 Get started with Arize in minutes using the new Arize skills with your coding agent.🧠 Learn more here: https://arize.com/docs/ax/agents/arize-skillsSkill...
002
Jim Bennett @jimbobbennett.dev · 10/03/2026
2 days in at @arize.bsky.social and they are shipping amazeballs features. AI skills for Arize to instrument, debug, evaluate. Without leaving your editor or terminal.
010
Jim Bennett @jimbobbennett.dev · 17/02/2026
In 68 days, I'll be running the London Marathon for Sage House Tangmere. Read all about this amazing charity, and why I and others are running for them in their latest blog post. If you can, please support us! Please also repost for reach! www.dementiasupport.org.uk/post/london-...
dementiasupport.org.uk
London Marathon Runners 2026
The streets of the capital are set to come alive this spring as the 2026 London Marathon prepares to welcome tens of thousands of runners and an estimated million spectators for one of the world’s mos...
000
Jim Bennett @jimbobbennett.dev · 17/02/2026
The one day a year to sing this song! It’s pancake day! youtu.be/lS9J4vJ5K78?...
youtu.be
HD Pancake Day Song - Maid Marian and her Merry Men - with Pancake Recipe!
YouTube video by Prosser Family Adventures
010
Jim Bennett @jimbobbennett.dev · 26/01/2026
3 months till the 2026 @londonmarathon.bsky.social . I'm running for my Dad who has dementia, raising money for a dementia support charity that provides a safe place for folks with dementia and their families. Please sponsor me! 2026tcslondonmarathon.enthuse.com/pf/james-ben...
2026tcslondonmarathon.enthuse.com
Jim's Fundraising for Dementia Support from Sage House
Hi, I'm Jim - the best looking one in the profile picture. The second best looking one is my Dad, and he has dementia. I'm running to raise money for Sage House, a small, local dementia support charit
010
Jim Bennett @jimbobbennett.dev · 18/01/2026
In 99 days I’m running the 2026 TCS London Marathon to raise money for Sage House, a charity that helps folks with dementia like my Dad. Please sponsor me as I try to raise $10,000 for this charity! 2026tcslondonmarathon.enthuse.com/pf/james-ben...
090
Jim Bennett @jimbobbennett.dev · 05/12/2025
AI engineers, or folks who want to be AI engineers! Want to learn about eval engineering so you can build AI apps that work? I'm teaching a free 5-part course on this very topic! Starts next Tuesday at 9am PT. Sign up now! luma.com/6q19vpzb
luma.com
Eval Engineering for AI app developers - Lesson 1: Hello Evals! · Luma
Learn Eval Engineering in this free, 5-part, hands-on course. 90% of AI agents don't make it successfully to production. The biggest reason is the AI engineers…
011
Reposted by Jim Bennett
Pomerium @pomerium.io · 01/12/2025
Join @jimbobbennett.dev from Galileo and @nickyt.online as they dig into real-time guardrails for AI agents December 11th. 👀 www.youtube.com/watch?v=4cqR... #AI #AIGuardrails #AgenticAI
youtube.com
Real-Time Guardrails for AI Agents
Jim Bennett, principal developer advocate at Galileo, joins Nick Taylor to discuss real-time guardrails for AI to provide more boundary layers.
012
Reposted by Jim Bennett
Nick Taylor @nickyt.online · 24/11/2025
Just scheduled! Looking forward to hanging with @jimbobbennett.dev from Galileo in early December to dig into real-time guardrails for AI agents! 👀 www.youtube.com/watch?v=4cqR... #AI #AIGuardrails #AgenticAI
youtube.com
Real-Time Guardrails for AI Agents
YouTube video by Pomerium
041