Sign in

Jim Bennett

@jimbobbennett.dev
1.2K followers 842 following 972 posts

World's most energetic dev rel Microsoft MVP. 🌈ally. I ❤️ Star Wars Lego & 🐻‍❄️. Father, husband. He/him.

PostsRepliesMedia
Jim Bennett @jimbobbennett.dev · 25/09/2026
Grabs some ibuprofen before reading…
010
Jim Bennett @jimbobbennett.dev · 04/09/2026
I keep hearing folks say "delete all your skills" when a new model comes out. No - skills can be pipelines of work, not just "this is how you interact with a service". Audit, eval, don't just blindly delete. I wrote about this here: jimbobbennett.dev/blogs/skills...
jimbobbennett.dev
Don't delete your skills, audit them
Boris Cherny says to delete your CLAUDE.md, skills, and hooks every six months. That's good advice for one kind of skill and dangerous for another. Here's how to tell them apart, and why the skills th...
132
Jim Bennett @jimbobbennett.dev · 26/08/2026
Oh yeah! This is a LOT of fun!
021
Jim Bennett @jimbobbennett.dev · 13/08/2026
It’s true! No idea why we can’t just call is testing and be done with it.
020
Jim Bennett @jimbobbennett.dev · 07/08/2026
🎥 Watch: youtu.be/NB8Rexhq0zQ
000
Jim Bennett @jimbobbennett.dev · 07/08/2026
How do you trust agents before they hit production? At @arize.bsky.social's Observe 2026, Salesforce's Manjit Singh walked through evaluating agents across the lifecycle — single agents to orchestrators & handoffs. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 05/08/2026
🎥 Watch: youtu.be/5DdmGt1bXVQ
000
Jim Bennett @jimbobbennett.dev · 05/08/2026
"Everyone is an AI agent builder — if you let them." At @arize.bsky.social's Observe 2026, CrewAI shared enterprise lessons on where agent ROI shows up and how to scale building across an org. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 03/08/2026
🎥 Watch: youtu.be/vaLuK01UaTA
000
Jim Bennett @jimbobbennett.dev · 03/08/2026
As agents become digital coworkers, access control built for humans starts to break. At @arize.bsky.social's Observe 2026, WorkOS's Michael Grinich explored the identity & security model for autonomous agents. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 31/07/2026
🎥 Watch: youtu.be/m0mS7lLLDaw
000
Jim Bennett @jimbobbennett.dev · 31/07/2026
What does agent adoption actually look like in production? At @arize.bsky.social's Observe 2026, Mastra shared patterns from thousands of teams — what separates shipped agents from stuck prototypes. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 29/07/2026
🎥 Watch: youtu.be/ckiyS91FDpU
000
Jim Bennett @jimbobbennett.dev · 29/07/2026
From experimentation to production, agents need whole-lifecycle platforms. At @arize.bsky.social's Observe 2026, Microsoft's Sebastian demoed building, deploying, evaluating & governing agents with Microsoft Foundry. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 27/07/2026
🎥 Watch: youtu.be/6PxOlWKUDXo
000
Jim Bennett @jimbobbennett.dev · 27/07/2026
Scaling agents from prototype to production is an infrastructure problem. At @arize.bsky.social's Observe 2026, Anyscale's Robert Nishihara explained how Ray scales RL, inference & multimodal AI — and why RL is having a moment. Sketchnoted 👇 🔗 Video link in the comments.
220
Jim Bennett @jimbobbennett.dev · 24/07/2026
🎥 Watch: youtu.be/EEvloPIxsFU
000
Jim Bennett @jimbobbennett.dev · 24/07/2026
The hardest problems in AI aren't model problems anymore — they're evaluation problems. At @arize.bsky.social's Observe 2026, Hamel Husain argued agents are bringing the data scientist back to AI engineering. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 22/07/2026
🎥 Watch: youtu.be/IbnFWFvctU0
000
Jim Bennett @jimbobbennett.dev · 22/07/2026
What are the smartest AI investors seeing before everyone else? At @arize.bsky.social's Observe 2026, Jaya Gupta of Foundation Capital shared where VC is flowing across the AI stack. Sketchnoted the fireside 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 20/07/2026
🎥 Watch: youtu.be/XRV2BBNuOXI
010
Jim Bennett @jimbobbennett.dev · 20/07/2026
When a hallucination is a regulatory + financial risk, responsible AI gets real. At @arize.bsky.social's Observe 2026, BlackRock shared how it deploys AI to support pros managing trillions — with real guardrails & evaluation. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 17/07/2026
🎥 Watch: youtu.be/tSnlcSvpNlo
000
Jim Bennett @jimbobbennett.dev · 17/07/2026
With autonomous agents, observability shifts from "what happened" to "why did the agent do that." At @arize.bsky.social's Observe 2026, AWS's Nate Slater explored how agentic AI changes observability. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 15/07/2026
🎥 Watch: youtu.be/akX6H4ytGtg
010
Jim Bennett @jimbobbennett.dev · 15/07/2026
Two of the fastest-growing open-source agent projects, one conversation. At @arize.bsky.social's Observe 2026, OpenClaw & Nous Research debated where agent frameworks go next — memory, skill creation, long-term learning. Sketchnoted 👇 🔗 Video link in the comments.
120
Jim Bennett @jimbobbennett.dev · 13/07/2026
🎥 Watch: youtu.be/DVzZqIzlsRk
000
Jim Bennett @jimbobbennett.dev · 13/07/2026
"AI agents need specs, not prompts." At @arize.bsky.social's Observe 2026, George Zhang argued the engineer's real job is specifying the hill agents climb — tests, evals, rubrics, constraints. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 10/07/2026
🎥 Watch: youtu.be/Lsp5YZ9Jj3Q
000
Jim Bennett @jimbobbennett.dev · 10/07/2026
What happens when a leading AI coding company turns its product inward? At @arize.bsky.social's Observe 2026, Cursor shared how it uses agents, evals & agent-powered workflows to build Cursor itself. Sketchnoted 👇 🔗 Video link in the comments.
130
Jim Bennett @jimbobbennett.dev · 08/07/2026
🎥 Watch: youtu.be/Hvd6DYUQJ84
010
Jim Bennett @jimbobbennett.dev · 08/07/2026
"Kubernetes is not your sandbox." At @arize.bsky.social's Observe 2026, the Daytona team argued K8s wasn't built for agent workloads, and walked through what agent-native infrastructure actually needs. Sketchnoted 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 06/07/2026
🎥 Watch: youtu.be/ijlrwKiYYM4
010
Jim Bennett @jimbobbennett.dev · 06/07/2026
Building agents is harder than the demos make it look. At @arize.bsky.social's Observe 2026, Anthropic's Marius Buleandra shared why agent failures compound in production, how to design evals that catch them, and why human review still matters. Sketchnoted 👇 🔗 Video link in the comments.
110
Jim Bennett @jimbobbennett.dev · 03/07/2026
🎥 Watch: youtu.be/c1xPkDi-038
000
Jim Bennett @jimbobbennett.dev · 03/07/2026
How do you improve a product with hundreds of millions of users? At @arize.bsky.social's Observe 2026, OpenAI's Stuart Sy showed how ChatGPT turns fragmented 'vibes' into evidence + action. I sketchnoted the talk 👇 🔗 Video link in the comments.
100
Jim Bennett @jimbobbennett.dev · 01/07/2026
🎥 Watch: youtu.be/errTnC59gVM
000
Jim Bennett @jimbobbennett.dev · 01/07/2026
What's next: fleets of long-running cloud agents that observe, evaluate & improve themselves — humans shift from doing the work to directing it. New: Signal, Agent Studio, Fleet Observability, Harness-as-a-Judge & PXI in open-source Phoenix.
100
Jim Bennett @jimbobbennett.dev · 01/07/2026
Agents stopped being demos this year — they're shipping code, fixing bugs, and running real workflows. In @arize.bsky.social's Observe 2026 keynote, the founders lay out what's next. I sketchnoted the whole keynote 👇 🔗 Video link in the comments.
100
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 24/06/2026
A field guide to four of the latest from @jimbobbennett.dev, where each one leaks, and why no single pass rate was ever going to survive this. arize.com/blog/long-h...
011
Jim Bennett @jimbobbennett.dev · 15/06/2026
The recording of my #MSBuild session is now live, where you can learn how to understand and fix your agents using open source tools like Arize Phoenix. build.microsoft.com/en-US/sessio... 2/2
build.microsoft.com
Understand and fix Agent Framework apps with observability and evals
Your AI apps are getting more complex, with multiple agents, tools, and different orchestration patterns. This makes them harder to understand, debug, and test. This session shows you how to visualize...
020
Jim Bennett @jimbobbennett.dev · 15/06/2026
Do you have an AI agent? Do you actually know what it is doing? Do you know if it works? Typically the answer to the first question is yes, and for the second it's we think so, based off 'vibes'. Which is a terrible way to build and run production software. 1/2
110
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 12/06/2026
Apple paid Google ~$1B/yr to license memory for Siri. OpenAI rebuilt ChatGPT memory in place. Anthropic gave models an API to consolidate their own. All called "memory." None is what users mean. @jimbobbennett.dev wrote a field map: arize.com/blog/memory...
arize.com
Memory is still a missing primitive: Cataloguing what the field is actually shipping
This week the field shipped four kinds of memory, and Apple paid Google a billion dollars a year for one of them. None of the four is what the demos imply. A field map of what's actually shipping, and the missing primitive that sits between the buckets.
012
Jim Bennett @jimbobbennett.dev · 12/06/2026
My slides have been the basic Keynote black theme, text on the left, image on the right, for years. Will never change so I can always point back to the before times and say it was always like this.
010
Jim Bennett @jimbobbennett.dev · 08/06/2026
New post on: → Why "I don't care" is an opinion (and a useful one) → Why the cliché needs to die → Why my C# opinion is built on a foundation, but my framework opinion is built on a Tuesday → The most important scene in the original Star Wars trilogy (fight me) www.linkedin.com/pulse/strong...
linkedin.com
Strong opinions, strongly held - and why I don't care about your tooling debate
I was in a meeting last week where the team was debating which of two tools to use for a job. Both of them do the thing.
000
Jim Bennett @jimbobbennett.dev · 08/06/2026
My version is: strong opinions strongly held, loose opinions loosely held. If I formed an opinion on data and 20 years of scars, it shouldn't flip on a clever argument over coffee. And if I haven't done the work to form one, "I don't care, just pick one" is a perfectly honest answer.
linkedin.com
Strong opinions, strongly held - and why I don't care about your tooling debate
I was in a meeting last week where the team was debating which of two tools to use for a job. Both of them do the thing.
100
Jim Bennett @jimbobbennett.dev · 08/06/2026
"I genuinely don't care. Pick one." That was my contribution to a meeting last week where the team was debating two tools. And it was the most useful thing I said all day. "Strong opinions, loosely held" is the "approved" take. I think it's mostly nonsense.
linkedin.com
Strong opinions, strongly held - and why I don't care about your tooling debate
I was in a meeting last week where the team was debating which of two tools to use for a job. Both of them do the thing.
110
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 03/06/2026
Microsoft picked OpenInference. Twice. The open trust stack for AI agents announced at #MSBuild, ASSERT for evaluation, ACS for controls, both ride on the open tracing standard Arize built for agents. arize.com/blog/micros...
arize.com
Microsoft's open trust stack runs on OpenInference
Microsoft's open trust stack for AI agents puts ASSERT and Agent Control Specification on top of OpenInference, connecting evaluation, runtime controls, and observability through a shared trace contract.
011
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 02/06/2026
At Microsoft Build? Our 2 must do things for today: 1. Catch Sarah Bird's session - Observe and control agents with OSS tools build.microsoft.com/en-US/sessi... 2. Head to the Microsoft AI expert booth to meet with @jimbobbennett.dev from our devrel team about AI observability and Evals #MSBuild
build.microsoft.com
Observe and control agents across any framework with open source tools
As AI agents move into production, developers own safety, governance, and reliability across Microsoft Agent Framework and open-source stacks. This session shows how to govern agents end to end: turning your requirements into context-aware evaluations, stress-testing against adversarial risks, applying open controls that work across frameworks, and keeping humans in the loop on high-stakes actions. Leave with a blueprint for shipping agents at enterprise scale. Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
011
Reposted by Jim Bennett
Arize AI @arize.bsky.social · 01/06/2026
Will you be at Microsoft Build this week, either in person in SF or virtually? Our very own @jimbobbennett.dev will be giving a demo session on understanding and fixing agents with open source observability and evals, Wednesday 3:30pm, Theater C. #MSBuild build.microsoft.com/en-US/sessi...
build.microsoft.com
Understand and fix Agent Framework apps with observability and evals
Your AI apps are getting more complex, with multiple agents, tools, and different orchestration patterns. This makes them harder to understand, debug, and test. This session shows you how to visualize the decisions your LLMs are making in complex Microsoft Agent Framework applications, using open standards and open source tooling to provide you with instrumentation and observability. You'll also see how you can use another LLM as a judge to evaluate how well your AI app is working. Seating for this session is first-come, first-served. Add it to your schedule to plan your day and arrive early to secure a spot.
011