Andrew Stellman 👾 @andrewstellman.bsky.social · 14hOne of the biggest challenges for new and intermediate developers trying to integrate AI into their learning is that an overreliance on AI-generated code can actually prevent them from learning. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 18hAs I’ve trained developers and teams who are increasingly adopting AI coding tools, I’ve noticed that the developers who adapt best aren’t always the ones with the deepest expertise in a specific framework. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 23hDevelopers are doing incredible things with AI. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 02/10/2026Many team leads, managers, and instructors looking to help developers ramp up on AI tools assume the biggest challenge is learning to write better prompts or picking the right AI tool; that assumption misses the point. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 02/10/2026Encourage pairing or small-group prompt reviews: Make AI-assisted development collaborative, not siloed. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 02/10/2026In the past, experienced developers could build deep expertise in a single technology (like Rails or React, for example) and that expertise would consistently get them recognition on their team and help them stand out in reviews and job interviews. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 01/10/2026Real work requires chaining multiple steps together into a pipeline, or a reproducible series of steps (some deterministic, some requiring an LLM) that lead to a single result, where each step’s output feeds the next. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 01/10/2026User stories intentionally focused the engineering work back on people and what’s in their heads. Whether it’s a requirements document in Word or a user story in Jira, the most important thing isn’t the piece of paper, ticket, or document we wrote. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 01/10/2026When you ask the AI to help write unit tests for generated code, first have it generate a plan for the tests it’s going to write. Watch for signs of trouble: lots of mocking, complex setup, too many dependencies—especially needing to modify other parts of the code. 210
Andrew Stellman 👾 @andrewstellman.bsky.social · 30/09/2026In AI-driven development, developers don’t just accept suggestions or hope the output is correct. They assign specific roles to specific tools: 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 30/09/2026Skipping the fundamentals: Using AI for extensive code generation before understanding basic software development concepts and patterns. Ensure learners can solve simple development problems on their own (without the help of AI) before accelerating with AI on more complex ones. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 30/09/2026Good framing starts by getting clear about the nature of the problem you’re solving. What exactly are you asking the model to generate? What information does it need to do that? Are you solving the right problem in the first place? 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 29/09/2026Vibe coding is a normal and useful way to explore with AI, but on its own it presents a significant risk. The models used by LLMs can hallucinate and produce made-up answers—for example, generating code that calls APIs or methods that don’t even exist. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 29/09/2026Try generating multiple solutions. Asking the AI to produce two or three alternatives forces it to vary its approach, which often reveals different assumptions or trade-offs. One version may be more concise; another more idiomatic; a third more explicit. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 29/09/2026I’ve spent over 20 years writing about software development for practitioners, covering everything from coding and architecture to project management and team dynamics. 110
Andrew Stellman 👾 @andrewstellman.bsky.social · 28/09/2026History keeps repeating itself: From binders full of scattered requirements to IEEE standards to user stories to today’s prompts, the discipline is the same. We succeed when we treat it as real engineering. Prompt engineering is the next step in the evolution of requirements engineering. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 28/09/2026Context dumping: Pasting entire codebases into prompts. Teach scoping—What’s the minimum context needed for this specific problem? Help them anticipate what the AI needs, and provide the minimal context required to solve each problem. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 28/09/2026One of the experiments I’ve been running as part of my work on agentic engineering and AI-driven development is a blackjack simulation where an LLM plays hundreds of hands against blackjack strategies written in plain English. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 27/09/2026Replacing the LLM validator with deterministic code (48% → 79%). This was the single biggest improvement in the entire arc. The pipeline had a second LLM call that scored how accurately the player followed strategy, and it was wrong 73% of the time. It applied its own blackjack intuitions 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 27/09/2026If you’re building anything where LLM output feeds into the next step, the same question applies to every step in your chain: Does this actually require judgment, or is it deterministic work that ended up in the LLM because the LLM can do it? The strategy validator felt like a judgment call 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 27/09/2026Every kickoff opens with what prior phases accomplished, includes explicit boundaries about what’s frozen, and names which future phase owns each piece of remaining work, because without it the AI will helpfully start doing Phase 3 work while you’re still in Phase 2. Each phase also ends with a 110
Andrew Stellman 👾 @andrewstellman.bsky.social · 26/09/2026Running pipelines at scale made the failures obvious and immediate, which, for me, really underscored an effective approach to minimizing the cascading failure problem: make deterministic work deterministic. That means asking whether every step in the pipeline actually needs to be an LLM call. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 26/09/2026You can start building a toolkit at any point in your project. The way it happened for me was organic: After weeks of working with Claude and Gemini on Octobatch configuration, the knowledge about what worked and what didn’t was scattered across dozens of chat sessions and context files. I 200
Andrew Stellman 👾 @andrewstellman.bsky.social · 26/09/2026I was working in Copilot with a seven-step plan, going through it one step at a time, having another AI review each step before moving on. Steps one and two went fine. When it came time to do step three and I gave it the prompt, it jumped straight to step four. This kind of thing can be really 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 25/09/2026The cost of poor context management is actually measurable. A developer on Microsoft’s Dev Blog recently timed his own reorientation overhead and found he was spending over an hour a day just reexplaining things to his AI that it had known in a previous session. He’s not alone. There are now 200
Andrew Stellman 👾 @andrewstellman.bsky.social · 25/09/2026I didn’t want to hang all this on one chat, so I went back and ran the same kind of self-examination on a handful of my other chats, doing completely different work: planning a course, writing up a guide, a couple of unrelated coding projects. The same pushiness showed up in every one. It 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 25/09/2026The right level is to state the principle, give one concrete example, and trust the AI to apply it to new situations. If you keep tightening guardrails every time an AI makes a judgment call, the signal gets lost in the noise and performance gets worse, not better. When something goes wrong, 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 25/09/2026Overriding the model’s priors (81% → 84%). One strategy required hitting on 18 against a high dealer card, which any conventional blackjack wisdom says is terrible. The LLM refused to do it. Restating the rule didn’t help. Explaining why the counterintuitive rule exists did: The prompt had to 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 24/09/2026The playbook was writing everything down to files as it went, which is why those runs could last that long at all. But I didn’t want that behavior. Running 15 million tokens in a single session is expensive, and if you’re on pay-as-you-go API tokens instead of a flat-rate plan like Copilot or 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 24/09/2026I’m not proposing a standard format for a toolkit file, and I think trying to create one would be counterproductive. Configuration formats vary wildly from tool to tool—that’s the whole problem we’re trying to solve—and a toolkit file that describes your project’s building blocks is going to 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 23/09/2026I’ve been working on the Quality Playbook, my open source AI skill that uses quality engineering to find bugs that normal AI code review misses, and I recently had a batch of work that turned into a long run of point releases. I was using Claude Cowork as the orchestrator: planning scope, 121
Andrew Stellman 👾 @andrewstellman.bsky.social · 23/09/2026The longer a defect goes uncorrected, the more entrenched it becomes and the more things get built on top of it. Context drift works the same way. When the AI loses track of a design decision early in a session, everything built on that lost context compounds the error. And just like a 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 23/09/2026“Do not paraphrase from memory” is the line that did the actual work. The instruction couldn’t trust the agent’s memory of what BUGS.md said, even though BUGS.md was sitting right there in the context window. So the instruction forced a fresh read of the file at the moment of writing. The 110
Andrew Stellman 👾 @andrewstellman.bsky.social · 22/09/2026What stays the same is that somebody owns the result. When I vibe-coded my bus tracker, nobody was going to catch that wrong stop ID but me. When Cherny directs tens of thousands of agents, nobody owns what they ship but him. The verification changes with the stakes. A throwaway prototype gets 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 22/09/2026I think the through line through all of this is that developers both overestimate and underestimate AI. We overestimate how much it can hold in its memory and its ability to remember things and make decisions for us. So we’ll just stuff a whole bunch of stuff in the context window and assume 110
Andrew Stellman 👾 @andrewstellman.bsky.social · 22/09/2026I was reviewing some output from Opus with GPT Astra, and it kept finding things Opus was telling it that were stupid, so I told it this. 000
Andrew Stellman 👾 @andrewstellman.bsky.social · 22/09/2026Early runs of my simulation had a 37% pass rate. The LLM would add up card totals wrong, skip the dealer’s turn entirely, or ignore the strategy it was supposed to follow. The big problem was that these errors compounded: If the model miscounted the player’s total on the third card, every 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 21/09/2026We’ve got a really good real-world example of how this can work. O’Reilly’s learning platform has an AI engine that answers questions out of the books on the platform, tells you which ones it drew on, and pays the authors and publishers behind them. It works because the corpus is small and 110
Andrew Stellman 👾 @andrewstellman.bsky.social · 21/09/2026The key to not getting frustrated when the AI loses track of steps or can’t seem to count from prompt to prompt is to remember what it’s good at and how it remembers things. If the AI you’re using does that, check the conversation history. You’ll probably see something like “summarizing 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 21/09/2026There’s a useful way to think about reliability problems like that: the March of Nines. Getting an LLM-based system to 90% reliability is the first nine, and it’s the “easy” one. Getting from 90% to 99% takes roughly the same amount of engineering effort. So does getting from 99% to 99.9%. Each 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 20/09/2026The U-shape says the model attends best to the beginning and end of its context. The natural move is to put your most load-bearing information in those positions and keep the middle for things you don’t need the model to focus on. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 20/09/2026From the agent’s perspective, compaction is seamless. It’s tracking state, referencing decisions made earlier in the conversation, and then at some point the earlier context is gone. But the agent can’t tell the difference between “I never knew that” and “I knew it but lost it.” 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 20/09/2026When you give an AI a multistep task, the natural instinct is to spell out the steps. First do this, then do that, then combine the results. The problem is that step-by-step procedures are the first thing the AI forgets when the context window fills up. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 19/09/2026On the other hand, we massively underestimate its ability as an orchestrator. Your prompt doesn’t just have to ask a question or ask the AI to generate something. 110
Andrew Stellman 👾 @andrewstellman.bsky.social · 19/09/2026The rule gives the model permission to be done. It makes stopping, with nothing queued, a legitimate way to finish a turn rather than something the model treats as leaving the job half-done. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 19/09/2026You see attribution everywhere once you start noticing it. Bibliographies and references are attribution. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 18/09/2026Look for the negative requirements. What should your software not do? What states should be impossible? What data should never be exposed? These negative requirements are often the most valuable because they define boundaries that structural review can’t see. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 18/09/2026I think loop engineering is a good name and an accurate one. Designing the loop that drives the agent is a real skill, and we need a word for it. 100
Andrew Stellman 👾 @andrewstellman.bsky.social · 18/09/2026Don’t run one long session. Run many short ones, each reading fresh from disk. 110
Andrew Stellman 👾 @andrewstellman.bsky.social · 17/09/2026We’re at the 640K stage of AI development. The context window is the new RAM ceiling. 100