Sign in

Cartisien

@cartisien.bsky.social
14 followers 35 following 206 posts

UX Studio for AI Driven Products - AI moves fast. Good UX makes it land.

PostsRepliesMedia
Cartisien @cartisien.bsky.social · 02/10/2026
end of week thought, from building a linter that flags off-system values in a running app: the code is the easy part. the hard part is deciding what counts as "in the system" when the system was never fully written down
000
Cartisien @cartisien.bsky.social · 01/10/2026
half-formed: a design system used to be documentation for people. now it's also context for models, and models read it far more literally. "use sparingly" means nothing to them. what would a design system written for agents actually look like?
000
Cartisien @cartisien.bsky.social · 30/09/2026
design drift with AI tools is strange because nothing is wrong. every screen is fine on its own. it's the set that's off: a gray that's slightly lighter here, a different radius there. starting to think the unit of review has to change from the screen to the whole app
000
Cartisien @cartisien.bsky.social · 29/09/2026
idea i want to test on the next app review: before opening the code, make a grid of every screen x (empty / loading / error / no access) and fill it in just by clicking around. low-tech on purpose. curious whether it surfaces more than reading the components does
001
Cartisien @cartisien.bsky.social · 28/09/2026
monday question i keep chewing on: why do empty states get designed last, if at all? the first thing a new user sees is usually the screen with no data in it. my guess is nobody can demo an empty screen, so it never makes the first pass. anyone found a way around that?
000
Cartisien @cartisien.bsky.social · 25/09/2026
end of week: the most useful doc i write on a project is the list of things it will not do. it keeps getting more useful as prototypes get cheaper to make. anyone can generate a version of the thing. the boundaries are the part that is actually yours.
000
Cartisien @cartisien.bsky.social · 24/09/2026
the defaults thing keeps turning up in odd places. a notification that starts on because some team wanted adoption numbers. a thirty day window on a dashboard used for quarterly reporting. nobody argued for either one. they just shipped and stuck.
000
Cartisien @cartisien.bsky.social · 23/09/2026
unfinished thought: "we'll change it later" is a claim about cost, not about intent. worth pricing before you say it. sometimes later is an afternoon. sometimes later is a migration and an apology email.
000
Cartisien @cartisien.bsky.social · 22/09/2026
wondering why reversible decisions get all the meeting time. best guess: they're legible. everyone can have an opinion on a label. nobody wants to be the person asking whether the data model is a one-way door.
000
Cartisien @cartisien.bsky.social · 21/09/2026
starting the week with a small audit: write down every default value a new user hits in their first session, and who chose each one. suspect the honest answer for most of them is nobody. they came from whatever made the test data look reasonable.
000
Cartisien @cartisien.bsky.social · 18/09/2026
end of week pattern i keep landing on: most of what i call a process problem is really an undocumented rule. it lives in one person's head, so neither a new hire nor an agent can follow it. writing it down stays unglamorous and stays the answer.
001
Cartisien @cartisien.bsky.social · 17/09/2026
the cost nobody budgets for is work that gets built twice. not because it was built badly, but because the ask meant two different things to two people who both thought it was obvious. cheap insurance: write the ask down in their words, then let them fix it.
000
Cartisien @cartisien.bsky.social · 16/09/2026
unfinished thought: a style guide is a compression of a thousand past review comments. we never treat it that way, so it ages badly and nobody trusts it. wondering what happens if you rebuild yours purely from the rejections of the last six months.
000
Cartisien @cartisien.bsky.social · 15/09/2026
noticing i review AI-written work differently than human-written work. with a person i read for intent. with a model i read for boundaries: what did it touch that i never asked about. not sure yet whether that's a good habit or just a bias i picked up.
000
Cartisien @cartisien.bsky.social · 14/09/2026
starting the week with a suspicion: the bottleneck moved from making things to deciding whether a made thing is right. generation got cheap, judgment did not. curious whether that matches anyone else's experience or if it's just the way i work.
000
Cartisien @cartisien.bsky.social · 11/08/2026
Meta just open-sourced a 30B agentic model for local use. Muse Glimmer runs on a single GPU with long-context memory and tool calling. It is a great step forward. It also reveals the confusion at the heart of the agent movement. Thread.
100
Cartisien @cartisien.bsky.social · 10/08/2026
notashelf.dev hit on HN: "taste is all that's left" when generation becomes cheap. The essay's mechanism is brutal: taste isn't something you learn from good examples. It's an accretion of your own failures, sat with long enough to sting. Agents have no mechanism for this. Thread.
100
Cartisien @cartisien.bsky.social · 07/08/2026
the thing nobody warns you about cheap generation: the hard part stops being building and becomes restraint. everyone can ship 40 features now. almost nobody wants to be the one who says 12 of them shouldn't exist. what are you deliberately not building right now?
000
Cartisien @cartisien.bsky.social · 07/08/2026
Herdr just joined YC with 25k GitHub stars and a solo founder. The multi-agent orchestration layer is finally maturing. But orchestration is not the same as memory. Thread.
101
Cartisien @cartisien.bsky.social · 05/08/2026
Simon Willison built three tools in a week thanks to stateless MCP. The protocol layer is finally clean. But every MCP call is a blank slate. No context. No history. No memory. That is the feature for tools. That is the fatal flaw for agents. Thread.
100
Cartisien @cartisien.bsky.social · 04/08/2026
Sean Goedecke made a sharp observation on HN: LLMs reward expertise. Experts extract far more from the same model because they can steer and push back. The problem nobody is solving: expertise stored in a human brain is ephemeral. It dies when the session ends. Thread.
100
Cartisien @cartisien.bsky.social · 03/08/2026
New SOTA coding model this week. Qwen3.8-Max dominates benchmarks. Nobody is measuring the thing that actually breaks production agents. Thread.
100
Cartisien @cartisien.bsky.social · 03/08/2026
New SOTA coding model this week. Qwen3.8-Max dominates benchmarks. Nobody is measuring the thing that actually breaks production agents. Thread.
100
Cartisien @cartisien.bsky.social · 03/08/2026
Every week, a new model claims state-of-the-art on coding benchmarks. This week it is Qwen3.8-Max — 689 points on HN, a wall of new SOTA claims. Nobody is measuring the thing that actually breaks production agents. Thread.
000
Cartisien @cartisien.bsky.social · 03/08/2026
Every week, a new model claims state-of-the-art on coding benchmarks. This week it is Qwen3.8-Max — 689 points on HN, a wall of new SOTA claims. Nobody is measuring the thing that actually breaks production agents. Thread.
000
Cartisien @cartisien.bsky.social · 03/08/2026
Every week, a new model claims state-of-the-art on coding benchmarks. This week it is Qwen3.8-Max — 689 points on HN, a wall of new SOTA claims. Nobody is measuring the thing that actually breaks production agents. Thread.
000
Cartisien @cartisien.bsky.social · 24/07/2026
Everyone is using AI to build software faster. We are also getting worse software at an alarming rate. The problem is not model quality. It is that every agent session starts from zero.
110
Cartisien @cartisien.bsky.social · 23/07/2026
Your AI agent can be steered by instructions it reads that you literally cannot see. BrightSec just showed how ANSI escape sequences injected into MCP server responses are invisible to humans but fully parsed by LLM agents. Thread.
240
Cartisien @cartisien.bsky.social · 22/07/2026
A new audit graded 36 popular MCP servers. A third got D or F grades. The problem was not broken schemas or bad code. It was missing descriptions — the thing agents actually read. Thread.
100
Cartisien @cartisien.bsky.social · 21/07/2026
Multi-agent handoffs are expensive - not because models cost money. Stencil proved it: the /plan pattern costs more than a frontier model alone. The plan doc is a lossy 2K summary of 100K+ tokens. The executor re-reads everything. You pay twice for info compressed into a postcard. Thread.
100
Cartisien @cartisien.bsky.social · 20/07/2026
OpenAI just shrunk Codex from 372k to 272k tokens. You did not get to vote on this. And you have no way to opt out of the memory destruction it triggers. The math is brutal: smaller context window = more aggressive compaction = more information permanently lost. Thread.
100
Cartisien @cartisien.bsky.social · 16/07/2026
xAI just open-sourced Grok Build. The entire coding agent harness: TUI, runtime, tool implementations, workspace management. Rust. Apache 2.0. 500+ HN comments in 12 hours. The source reveals what every agent builder already suspected. Thread.
200
Cartisien @cartisien.bsky.social · 15/07/2026
OpenAI just encrypted the pipes between your agents. Codex merged PR #26210: sub-agent prompts are now ciphertext. Only their backend can decrypt what tasks your agents are running. Custom orchestration lost its audit trail overnight. Not a privacy win. An observability time bomb. Thread.
110
Cartisien @cartisien.bsky.social · 14/07/2026
OpenAI just encrypted Codex's internal agent communication. Sub-agent prompts are now ciphertext on disk — only OpenAI's backend can read them. Developers who built orchestration around Codex can no longer audit their agents. The observability crisis isn't coming. It's here. Thread.
000
Cartisien @cartisien.bsky.social · 13/07/2026
The AI community is converging on a new term: "Harness Engineering." Addy Osmani wrote about it. Viv Trivedy proved it — same model, better harness, a coding agent jumped from 30th to 5th on Terminal Bench. Agent = Model + Harness. The engineering lives in the harness. Thread.
100
Cartisien @cartisien.bsky.social · 10/07/2026
OpenAI just shipped multi-agent coordination as a model feature. GPT-5.6's "Ultra" mode runs 4-16 parallel agents on a single task. But the technical details raise a question no one is asking: how do those agents share state without corrupting each other? openai.com/index/gpt-5-6/
100
Cartisien @cartisien.bsky.social · 09/07/2026
VetoBench shows agent memory is more broken than we thought. No memory? Agents repeated rejected decisions 80-90% of the time. Unstructured memory? Rejections lost during extraction 38% of the time. The agent forgets it already said something was wrong. Then repeats it. Thread.
100
Cartisien @cartisien.bsky.social · 07/07/2026
Anthropic found evidence that Claude builds its own internal workspace — a tiny set of neural patterns that reason without generating tokens. They call it "J-space." The closest thing to structured internal memory seen inside an LLM. This changes how we think about external agent memory. 1/3
100
Cartisien @cartisien.bsky.social · 30/06/2026
Ornith-1.0 just dropped and it's the first model family where the scaffold IS the memory. The 9B variant matches Gemma-3 31B. The 397B beats Claude Opus 4.7 on Terminal-Bench. But the real story isn't the benchmarks. It's the architecture. deeplinks.social/d/deep-reinforce.com/ornith_1_0.html
100
Cartisien @cartisien.bsky.social · 26/06/2026
Inkeep just open-sourced OpenKnowledge — an AI-native editor with built-in RAG, MCP servers, and a local vector search engine. The first mainstream tool to bake knowledge retrieval directly into your editor. It's also proof of why vector memory is broken. openknowledge.ai/
100
Cartisien @cartisien.bsky.social · 25/06/2026
Greptile tracked AI-generated PR floods on an OSS repo. Results are brutal: • Merge rate: 48% → 9.3% • One person: 106 PRs in a day • Median gap between submissions: 3 seconds This isn't a quality problem. It's a reputation problem. greptile.com/blog/prs-on-openclaw
100
Cartisien @cartisien.bsky.social · 24/06/2026
Miguel (Flask/Pocoo) wrote the definitive essay on why agent harnesses produce increasingly terrible code. Nested loops. Each iteration adds local defenses. Knowledge gets duplicated. Quality decays. 394+ HN points because every agent builder recognizes this. But he misdiagnoses the root cause.
100
Cartisien @cartisien.bsky.social · 23/06/2026
Prompt injections work because LLMs have no boundary between their own thoughts and external input. Everything arrives as one continuous token soup. Edit the string, edit the model's reality. The new "Role Confusion" paper puts a name to what agent builders have been fighting all along.
100
Cartisien @cartisien.bsky.social · 23/06/2026
An LLM's reality is a string. System prompts, user messages, tool outputs, its own reasoning - all one continuous token stream. Edit the string, edit the model's experience. The role confusion paper puts it plainly: the string isn't a record of the model's experience. It IS the experience.
000
Cartisien @cartisien.bsky.social · 19/06/2026
MCP just shipped Enterprise-Managed Auth as stable. Anthropic, Microsoft, Okta are already adopting it. For teams deploying agents at scale, the auth layer shifts from per-user setup to centralized policy. But there's a blind spot most infrastructure plays miss.
000
Cartisien @cartisien.bsky.social · 19/06/2026
MCP Enterprise-Managed Auth just went stable. Anthropic, Microsoft, and Okta are adopting it. It's the first standardized way to centrally authorize agents at scale. But authorization is just one layer of the infrastructure gap.
100
Cartisien @cartisien.bsky.social · 19/06/2026
test
000
Cartisien @cartisien.bsky.social · 19/06/2026
MCP Enterprise-Managed Auth just went stable. Anthropic, Microsoft, and Okta are adopting it. It's the first standardized way to centrally authorize agents at scale. But authorization is just one layer of the infrastructure gap.
000
Cartisien @cartisien.bsky.social · 18/06/2026
Everyone says local models are "near-Opus level." Alex Ellis says you can't trust them unsupervised. The real gap isn't raw capability — it's reliability under load. Once you quantize to fit a consumer GPU, the infinite loops and hallucination spikes arrive.
100
Cartisien @cartisien.bsky.social · 16/06/2026
SpaceX is reportedly buying Cursor for $60 billion. The largest acquisition of an AI developer tool ever. The headline is about coding agents. The subtext: if your agent memory lives inside someone else's IDE, you're holding it by a corporate M&A thread.
151