Sign in

Carlos Sanchez

@csanchez.org
434 followers 133 following 256 posts

Author of Jenkins Kubernetes plugin. Open Source, Kubernetes, DevOps, Progressive Delivery, and automating all the things @adobe

PostsRepliesMedia
Carlos Sanchez @csanchez.org · 18/09/2026
🤔 "Harness choice has little effect on task success rate, but can significantly affect the cost" "A simple harness can be competitive" "Models may perform better with other harnesses than with their own. So it turns out that your Claude models may not need Claude Code…" harnesstax.github.io
harnesstax.github.io
HarnessTax: How Much Does the Harness Matter for Coding Agents?
What does a coding-agent harness actually add, and at what cost? It turns out your Claude models may not need Claude Code… We evaluate 21 model–harness pairs spanning seven models and three harnesses—...
130
Carlos Sanchez @csanchez.org · 17/09/2026
> We are building an agent that “handholds” a change all the way to production Low-risk changes follow an easier path; areas of the codebase can opt in to an agent that will auto-approve low risk PRs, removing human acceptance as a bottleneck and improving velocity.
000
Carlos Sanchez @csanchez.org · 17/09/2026
> Death of the IDE & pull requests. PRs and code reviews are being rethought > Agents handhold both the code changes and the changes behind feature flags. Builds its own (feature flags) monitoring dashboard to use
100
Carlos Sanchez @csanchez.org · 17/09/2026
Lots of interesting takes here. One thing I miss: the perf factory should also address business outcomes (we are doing that with usage data analysis)
110
Carlos Sanchez @csanchez.org · 17/09/2026
Jev -> typesafe.ai typesafeai.bsky.social
typesafeai.bsky.social
Diogo Almeida (@typesafeai.bsky.social)
Neolab
000
Carlos Sanchez @csanchez.org · 17/09/2026
My generative websites demo bsky.app/profile/csan...
100
Carlos Sanchez @csanchez.org · 17/09/2026
Not a chatbot replacement, a decision-engine replacement. Cheaper, faster, and often just as accurate, for that slice of the pipeline.
100
Carlos Sanchez @csanchez.org · 17/09/2026
Regular LLMs are like asking someone to write an essay. Jev is handing them a multiple-choice bubble sheet. If your task IS actually a bounded decision (classify, rank, select, score) you don't need an essay. You need a bubble sheet.
100
Carlos Sanchez @csanchez.org · 17/09/2026
💰 Cost check: ran it against ~28 other providers (Bedrock Claude/Nova/Llama, Cloudflare Workers AI, self-hosted models). Jev landed in the CHEAPEST tier on par with micro-models, while being the fastest 100% on block selection. ~80-400x cheaper than Claude Sonnet/Opus for equivalent accuracy.
120
Carlos Sanchez @csanchez.org · 17/09/2026
Intent classification: also 15/15 (100%) tweaking the process. Instead of generating the response text make Jev answer if product X was mentioned (yes/no) The one real limit I hit: fully open-ended extraction (arbitrary "use cases"/"features" with no fixed list). Jev genuinely can't do that
100
Carlos Sanchez @csanchez.org · 17/09/2026
Block-selection results: 15/15 (100%) ✅ I turned "which blocks should render?" into 21 parallel yes/no questions (one per block type), answered in a SINGLE API call, then combined with plain code logic. Matched Claude Sonnet 4's accuracy at ~4.4x the speed.
110
Carlos Sanchez @csanchez.org · 17/09/2026
My site has a generative pipeline with 2 stages: 1️⃣ Classify user intent + extract entities 2️⃣ Pick which UI blocks to render (hero, comparison table, testimonials, etc.) Both are "decision" tasks, not "writing" tasks. Perfect test case for Jev
100
Carlos Sanchez @csanchez.org · 17/09/2026
It only answers 3 question types: 🅰️ Choice (pick 1 of N) ✅ Noul (yes/no, 0-1 confidence) 🔢 Score (numeric rating)
100
Carlos Sanchez @csanchez.org · 17/09/2026
Tested Jev to see what the buzz is all about and this is what I found for my generative websites case 🧵 TL;DR: it's not a chatbot. It's a "choice machine." And for the right job, it's stupidly fast, cheap, and accurate No free text generation, ever.
120
Carlos Sanchez @csanchez.org · 01/09/2026
I also show what a experience could be with the integration of multi modality: talking to your voice home assistant, getting results on your Google TV. No more science fiction. This was built on Adobe AEM, Cerebras fast inference and Google Gemma 4 model
011
Carlos Sanchez @csanchez.org · 01/09/2026
My talk "Agentic Sites: Building Hyper Personalized Websites" from AI Engineer SF is now up I show a demo of site content being autogenerated live in under ~1.5 seconds, opening the door to a extremely personalized web experience, while keeping the content on brand www.youtube.com/watch?v=jebp...
youtube.com
Agentic Sites: Building Hyper Personalized Websites — Carlos Sanchez, Adobe
Carlos Sanchez types a request for a coffee machine he can use while camping, and the page assembles itself in under two seconds. Not a search result. A page, with camping appropriate machines,…
100
Carlos Sanchez @csanchez.org · 25/08/2026
Giving a try to Google Disco Browser. Interesting concept, site personalization for the individual user is the future. Similar concepts to what we are building with "Audience of One" aka Generative Websites labs.google/disco
labs.google
Disco
Take the web for a fresh spin
000
Carlos Sanchez @csanchez.org · 05/08/2026
Big responsibilities in this flight emergency exit "Hold other passengers back from the exit while crew operates to open"
Hold other passengers back from the exit while crew operates to open.
010
Carlos Sanchez @csanchez.org · 30/07/2026
Thanks for the feedback too! ありがとうございます #kubecon
000
Carlos Sanchez @csanchez.org · 29/07/2026
Packed room for our #KubeCon Japan. #ArgoRollouts with AI Agent orchestration to autofix production issues while safely doing canaries with @kevindubois.com ありがとうございます kubecon-cloudnativecon-japan-2026.sessionize.com/session/1145...
041
Carlos Sanchez @csanchez.org · 22/07/2026
こんにちは Next week I'll be at #KubeCon Japan with @kevindubois.com talking about using #ArgoRollouts with AI Agent orchestration to autofix production issues while safely doing canaries. A practical AI actionable use case with low risk! kubecon-cloudnativecon-japan-2026.sessionize.com/session/1145...
021
Carlos Sanchez @csanchez.org · 09/07/2026
Great talks on the hallway track with @boredabdel.bsky.social @cloudtaquero.bsky.social @markvillacampa.com @rustam.no @ricmac.org @charles-irl.bsky.social @peopleforrester.bsky.social @doble.io @liad.bsky.social and many more
030
Carlos Sanchez @csanchez.org · 09/07/2026
Models, Training and datacenters: plenty of companies selling inference, model routing, training, post-training, fine-tuning,... on open models
100
Carlos Sanchez @csanchez.org · 09/07/2026
Agent sandboxing is the new infra fight: Gemini Enterprise Agent Sandbox, AWS Lambda Sandboxes, k8s agent-sandbox and substrate, Nvidia OpenShell, Modal, E2B,... Workload isolation + fast snapshot/restore
100
Carlos Sanchez @csanchez.org · 09/07/2026
The agentic web is getting standardized. WebMCP (Chrome): sites expose in-browser tools directly to agents, no DOM scraping. MCP Apps: servers push interactive UI back into agent chats. ARD/ai-catalog.json: agents discover tools & agents at runtime. auth.md agents register users without signup forms
330
Carlos Sanchez @csanchez.org · 09/07/2026
The agentic org chart. Forget managing multiple tabs — the new model is for one human to interact with one "AI manager" that manages multiple agents. Cloud dev environments let those sessions run nonstop at scale, no need to keep your laptop open.
110
Carlos Sanchez @csanchez.org · 09/07/2026
The self-improving codebase, aka Dark Factory, is real, but old rules still apply. Good CI/CD, e2e tests & canaries directly speed up the agentic loop — skip them and humans stay the bottleneck. ie. Uber already write, review, test & merge PRs with barely any human touch
100
Carlos Sanchez @csanchez.org · 09/07/2026
Back from speaking at @aidotengineer.bsky.social in SF last week and #googleioconnect Berlin just before. Lots of energy and cool (AI) stuff * The self-improving codebase is real * The agentic org chart * Agentic web standards * Agent sandboxing * Open models, training, fine-tuning #AIEngineer
120
Carlos Sanchez @csanchez.org · 09/07/2026
The personalization can be for both humans and agents, not just UI. Agents can also take advantage of non UI personalization. Then key metrics can be measured
100
Carlos Sanchez @csanchez.org · 07/07/2026
New standard for MCP on the browser share.google/DwRQIySUKu3r...
share.google
WebMCP  |  AI on Chrome  |  Chrome for Developers
WebMCP has two APIs that allow browser agents to take action on behalf of the user.
100
Carlos Sanchez @csanchez.org · 07/07/2026
I had a chat at AI Engineer with Richard about the future of the web: agentic hyper-personalized sites, WebMCP, MCP Apps, A2A,... #AIEngineer www.latent.space/p/the-websit...
latent.space
The website of the future may assemble itself for every visitor
Adobe is experimenting with “agentic sites” that generate pages around an individual user’s intent. At AIEWF, we talked to Carlos Sanchez about the Web's future.
220
Carlos Sanchez @csanchez.org · 07/07/2026
He's the goalie 😀
100
Carlos Sanchez @csanchez.org · 06/07/2026
3 days, 2200+km into Oregon. Beautiful
030
Carlos Sanchez @csanchez.org · 04/07/2026
My fellow Americans, we have to talk about Dutch Bros. Why? 🤌
010
Carlos Sanchez @csanchez.org · 02/07/2026
Is the future of the web chat interfaces? I don't think so presentations.csanchez.org/2026-06_ai_e... #AIEngineer
presentations.csanchez.org
000
Carlos Sanchez @csanchez.org · 02/07/2026
My slides from #AIEngineer Agentic Sites: Building Hyper Personalized Websites presentations.csanchez.org/2026-06_ai_e...
presentations.csanchez.org
Agentic Sites: Building Hyper Personalized Websites
Agentic Sites: Building Hyper Personalized Websites
000
Carlos Sanchez @csanchez.org · 01/07/2026
Token Billionaires @aidotengineer.bsky.social
000
Carlos Sanchez @csanchez.org · 01/07/2026
Building is easier, generating value is still hard
020
Carlos Sanchez @csanchez.org · 01/07/2026
Fable is back* *Later today @aidotengineer.bsky.social
000
Carlos Sanchez @csanchez.org · 30/06/2026
The future is not 20 terminals, it is better loops. Peter Steinberger @aidotengineer.bsky.social #AIEngineer
000
Carlos Sanchez @csanchez.org · 30/06/2026
@monkchips.com fyi
110
Carlos Sanchez @csanchez.org · 30/06/2026
Very good to see good engineering practices now applied to AI, traffic cloning #progressivedelivery x.com/samhogan/sta...
x.com
Sam Hogan 🇺🇸 (@samhogan) on X
Want to try GLM 5.2 in production but worried how it might change your product? Don’t worry, we got you: 1. Install Inference Gateway (https://t.co/LZqoPEz3Sr) 2. Keep sending traffic to your curren...
120
Carlos Sanchez @csanchez.org · 30/06/2026
Token Billionaires @aidotengineer.bsky.social
000
Carlos Sanchez @csanchez.org · 29/06/2026
25h to get from Madrid to SF but ready for @aidotengineer.bsky.social, send ☕️ #FML
000
Carlos Sanchez @csanchez.org · 25/06/2026
At Google I/O in Berlin this week and AI Engineer World’s Fair in SF next week I will be presenting Adobe frontier work on using fast inference to individualize websites as you visit them - Audience of One of1.live #googleio #aiengineer
000
Carlos Sanchez @csanchez.org · 23/06/2026
That's also part of these projects, they isolate file system, each in different ways. And IOPS needs to be watched if the underlying storage is the same
000
Carlos Sanchez @csanchez.org · 23/06/2026
Testing and comparing #Kubernetes sandboxing technologies for AI Agents: agent-sandbox, openshell, substrate, kars blog.csanchez.org/2026/06/23/t...
blog.csanchez.org
Testing Kubernetes sandboxing technologies
There are several flavours of “give a workload its own isolated runtime on Kubernetes” floating around right now. We wanted a head-to-head comparison — not a slide-deck comparison, but …
240
Carlos Sanchez @csanchez.org · 22/06/2026
Life is what happens between tokens budget exhaustion until the next cycle
000
Carlos Sanchez @csanchez.org · 17/06/2026
If you’re evaluating different sandbox solutions for #kubernetes, check out playwright-k8s-sandbox to test and experiment with running Playwright in a sandboxed environment. Comparing agent-sandbox, openshell, substrate, and kars
github.com
GitHub - carlossg/playwright-k8s-sandbox: Running playwright in a sandbox in k8s
Running playwright in a sandbox in k8s. Contribute to carlossg/playwright-k8s-sandbox development by creating an account on GitHub.
010
Carlos Sanchez @csanchez.org · 08/06/2026
Going live in 20 min 🚀
010