Fred Hebert @ferd.ca · 27/09/2026In 1997, James Reason published 'Managing Organizational Accidents'. It was influential in popularizing his concepts around the Swiss Cheese model, his specific brand/recipe of safety (and just) culture, etc. and discussed ever since. But it also has this line in it just throwing incredible shade: 23011
Fred Hebert @ferd.ca · 26/09/2026This is a good example of the substitution myth in automation, applied to software workflows. Coactive work and the integration of various functions is different from the individual functions being doable in isolation. 0125
Fred Hebert @ferd.ca · 20/09/2026mornings are starting to be colder which means it is once again time to hover my hands above the toaster while it cooks toast like it’s a campfire 1271
Reposted by Fred HebertLorin Hochstein @norootcause.surfingcomplexity.com · 19/09/2026Today's blog post: a play in one act. surfingcomplexity.blog/2026/09/19/h...surfingcomplexity.blogHow the incident happened: a play in one actA hallway in a tech company office. A and B are standing in the hallway. B is holding a multiple page document. A: Wow, that incident from a couple of weeks ago was a real doozy. Do you know how it… 1114
Fred Hebert @ferd.ca · 06/09/2026I think it’s now been more than a year since I last wrote software just for fun (with or without AI). It feels more like chores or work by now. It feels like it had been a long time coming, but it just became clearer recently. 2180
Reposted by Fred Hebertresilienceinsoftware.org @resilienceinsoftware.org · 01/09/2026Another very cool post you might have missed from our community, on how expertise copes with overload, highlighted by treating incidents as first-class work: resilienceinsoftware.org/news/11453533resilienceinsoftware.orgExpertise and Overload | Resilience in Software FoundationResilience engineering views incidents through a different frame than the conventional approach in the software industry, which tends to treat incidents as an i 022
Fred Hebert @ferd.ca · 01/09/2026If I were a scientist my email signature would absolutely contain the words “more research is needed” 0121
Fred Hebert @ferd.ca · 27/08/2026I just realized I hadn’t shared this one yet, but I wrote about how the team I’m on restructured its workflows to lean into the code review bottleneck rather than trying to eliminate it when code generation took over, and what that ended up doing: www.honeycomb.io/blog/embraci...honeycomb.ioHow I Came to Embrace the Code Review BottleneckFaced with an endless stream of AI-generated code reviews, our team made the counterintuitive choice to lean into the bottleneck rather than reduce it. 3307
Reposted by Fred HebertLorin Hochstein @norootcause.surfingcomplexity.com · 26/08/2026Every risk register should have an entry that reads “we have misjudged the risks” which is in the “likelihood: high” and “impact: high” region of the risk matrix. 5378
Fred Hebert @ferd.ca · 20/08/2026Fix code review bottlenecks by doing like private torrent trackers and only allowing people to get their PR reviewed if they reviewed enough PRs beforehand to keep their ratio high enough. 1578
Fred Hebert @ferd.ca · 19/08/2026The first rule of the papers-reading club is that you’re unfortunately already in the papers-reading club and I will send you links and summaries 0140
Fred Hebert @ferd.ca · 17/08/2026Yet again hearing how some skilled work never truly mattered now that it's getting automated, while it absolutely did and was a point of professional pride for many. It's a significant aspect of 'deskilling', which is a known consequence of automation. Pretending otherwise is needlessly unkind. 1357
Fred Hebert @ferd.ca · 17/08/2026this seems more accurate given the frequency of errors right now 0121
Fred Hebert @ferd.ca · 14/08/2026Although there have been changes in cognition and systems theory, I wanted to bring up some stuff Rasmussen published in the 80s that was really elegant. I'll use 3 diagrams he published, covering the ideas of abstraction hierarchies and how people operate systems when troubleshooting them. 1136
Fred Hebert @ferd.ca · 10/08/2026This is a very good text on comparing ecology and organizational dynamics when it comes to harvesting signals; big fan of this one. psychsafety.com/a-practical-...psychsafety.comOrganisational Indicator Species: being an Organisational EcologistThe more elaborate and expensive our system for understanding something, the less we may actually understand it. Ecologists face the same problem — you can't measure the health of an ecosystem directl... 073
Fred Hebert @ferd.ca · 10/08/2026Wrote up a bunch of stuff about some patterns in system design, about the tension, contrast, and possibility of composing approaches of analytical decomposition to increase control, and of complexity-aware stances for emergence, and some pitfalls of either stance: ferd.ca/control-and-...ferd.caControl and complexity: tension in systems designcomposing two broad approaches, one based on analytical decomposition that aims to maintain control over a system, and one based on a perspective of complex systems that resist analysis, and implicati... 0177
Reposted by Fred HebertLorin Hochstein @norootcause.surfingcomplexity.com · 03/08/2026Quick little brainstorm-y blog post about traditional versus resilience engineering focuses (foci?): surfingcomplexity.blog/2026/08/02/t...surfingcomplexity.blogTraditional versus resilience engineering viewsAs a fan of resilience engineering, I often differ with people on where we should focus our scarce engineering cycles in order to improve reliability. I thought it would be a useful exercise to bra… 172
Fred Hebert @ferd.ca · 24/07/2026Delegating pressure to the final individual, who is now tasked with continuously fixing the entire system’s misalignments through their personal choices. 2254
Fred Hebert @ferd.ca · 14/07/2026I was a reviewer for this book so I can give you some spoilers about it: it rules, get a copy of it, start telling everyone about cumulative culture, unbuild the myths that keep your teams down, make shit be meaningful. 1288
Fred Hebert @ferd.ca · 13/07/2026One of the ironies about AI agents in ops tasks is that it feels like there has never been as much interest in creating a forgiving environment with proper structural support than through promising to remove humans from it, finally forcing a less individualistic and blameful approach to design. 7478
Fred Hebert @ferd.ca · 10/07/2026at this point why not just host my repos on the staging servers, github 081
Fred Hebert @ferd.ca · 03/07/2026These Angine de Poitrine fans got the best flags I’ve seen in a while; that’s some amazing concert gear. 0262
Fred Hebert @ferd.ca · 23/05/2026Even after years, one of the weirdest parts of gardening to me is needing to harden the seedlings before transplanting them. Like “yes hold on a minute I gotta take the plants out so they can play outdoors for a while, but they gotta be in before streetlights turn on” is a real and necessary thing. 0100
Fred Hebert @ferd.ca · 19/05/2026The ongoing stream of software engineering pieces that mention that the future is in writing spec but never bother to define what a specification is or at what abstraction levels it should be is appalling; arguably, tickets are a spec, the code is a spec, and work between both is connecting dots. 3315
Fred Hebert @ferd.ca · 17/05/2026“decisions with lasting social consequences are attached to a future moment when the technology is assumed to have revealed its true form, rather than addressed in the present in which it already operates. […] it defers responsibility by attaching accountability to a moment that never materializes”link.springer.comWaiting for AGI - AI & SOCIETYAI & SOCIETY - 03210
Reposted by Fred HebertLorin Hochstein @norootcause.surfingcomplexity.com · 10/05/2026New blog post about flipping the bozo bit: surfingcomplexity.blog/2026/05/09/f...surfingcomplexity.blogFlipping the bozo bit on flips the learning offI’m too young to have seen Bozo the Clown myself, but I’m old enough to get the references “Flipping the bozo bit” is an expression from the software world. Think about a ti… 1114
Fred Hebert @ferd.ca · 02/05/2026“[…] we reached for Recon, an amazing tool for diagnosing issues […] (the related Erlang in Anger is more-or-less required reading as all on-call engineers end up scouring its pages eventually)” Wild! I wrote these 10 years ago to help coworkers, and they’re still useful for real world issues now!discord.comYou’ve Got (Too Much) Mail: Behind the Scenes of the 3/25/26 Voice OutageOn March 25th, voice and video on Discord suffered major degradation beginning at 12:13 PDT, lasting a little over three hours. Learn how the issue originated, how it affected systems across Discord, ... 0192
Fred Hebert @ferd.ca · 29/04/2026Infinite love to my SRE coworkers, one of whom casually dropped this single line in a retro: "the industrial hourly deploy train and its consequences have been a disaster for society" 15717
Fred Hebert @ferd.ca · 20/04/2026Writing my action items on a piece of paper that I will burn to ward off spirits 092
Fred Hebert @ferd.ca · 18/04/2026I wrote for @resilienceinsoftware.org on "Superficial Blamelessness", where under the label of "blamelessness", we avoid punishing people, yet still focus fixes and interventions based on the same individualistic framing rather than a broader systemic stance. resilienceinsoftware.org/news/11502437resilienceinsoftware.orgSuperficial BlamelessnessIn 2012, John Allspaw (then CTO of Etsy) wrote a seminal blog post on the need for what he called Blameless Postmortems. Built off the notion of a “Just Culture” from the research of Sidney Dekker, he... 0207
Fred Hebert @ferd.ca · 08/04/2026I've gotten an early copy of Crisis Engineering by Marina Nitze, Matthew Weaver, and Mikey Dickerson, and just in time for its release today, here's my review of it: ferd.ca/notes/on-cri... TL:DR; I like it, good perspectives and an interesting mix of approaches. 092
Fred Hebert @ferd.ca · 07/04/2026One of my assignments asks whether a given set of psychological constructs amount to “good science” and it’s very weird as someone who’s written zero papers and done no real science to just attempt to go “sure dawg it’s okay science I guess but these hundreds of researchers could pick better models” 140
Fred Hebert @ferd.ca · 03/04/2026Every time I am faced with this Swedish login form, I have to say 'Logga in' out loud with a terminator voice. It is one of the few rules we can't change and must be respected. 2271
Fred Hebert @ferd.ca · 23/03/2026hell yeah @resilienceinsoftware.org swag is in! the law of requisite variety states that only variety in the regulator can destroy variety in the system being regulated. so if you need to deal with complexity you know you gotta join the club & begrudgingly increase complexity to keep things simple 1152
Fred Hebert @ferd.ca · 16/03/2026would you rather have many incidents of various small to moderate size for the foreseeable future or just one very big incident and then be done with them for good? 880
Fred Hebert @ferd.ca · 08/03/2026As usual @grimalkina.bsky.social is worth listening to. These principles are also worth considering and applying in all sorts of contexts. Here’s a sample from safety research I happened to read just yesterday (Dekker - Reconstructing human contributions to accidents, 2002) that aligns with it! 2295
Fred Hebert @ferd.ca · 07/03/2026When I joined the program I'm currently in, I told myself I would not read more papers and technical books outside of it because I'd need to balance about my energy levels—not spending it all on this. I was right, but also there's lots of other cool nerd shit I want to read through now and welp. 0120
Reposted by Fred HebertLiz Fong-Jones (方禮真) @lizthegrey.com · 04/03/2026Here is the fuller writeup I promised, a little bit overdue (I said Jan but it actually went out in Feb). Credit to @ferd.ca for the writeup, as well as to all of our incident responders who worked this 12+ day incident. www.honeycomb.io/blog/inciden...honeycomb.ioIncident Report: Exercises, Cleanups, and EvacuationsEvery year, Honeycomb runs disaster recovery scenarios in multiple environments, including in production. Although each of our instances runs in a single region, on at least three Availability Zones ... 0122
Fred Hebert @ferd.ca · 01/03/2026I really need to do this year’s garden planning. It’s gonna be time for seedlings in a few weeks and it’s gonna be nice to once again get going on one of them hobbies where you can’t really obsessively dictate the pace nor feel pressured in going faster. Just watch the plants grow and see what goes. 1150
Fred Hebert @ferd.ca · 26/02/2026Back in December, we had a large outage at work. The internal investigation took a while and the internal report was roughly 40 pages long. For the public, we managed to try and condense it to a much shorter format that we think can still offer useful insights to other organizations:honeycomb.ioIncident Report: Exercises, Cleanups, and EvacuationsEvery year, Honeycomb runs disaster recovery scenarios in multiple environments, including in production. Although each of our instances runs in a single region, on at least three Availability Zones ... 1124
Fred Hebert @ferd.ca · 25/02/2026I believe I had a relatively intuitive sense of how much Swiss cheese I could consume in one sitting before I can take no more. I also believe I now have a fairly empirical sense of how much literature about the Swiss Cheese Model I can consume in one sitting before I can take no more. 3170
Reposted by Fred HebertChastity Blackwell @blackisis.bsky.social · 23/02/2026I also think this discussion about how SRE work is being devalued by these products is at least parallel to the discussion about how people are willing to write clear documentation for *AI* consumption, but never put any value on it when it was for *actual people*. 082
Fred Hebert @ferd.ca · 23/02/2026AI SREs are framed as nameless job automation; Coding Assistants as named partners. The framing used to create these products reveals a lot about how the builders and buyers perceive these roles. I also write about the challenges and risks of picking self-limiting analogies in building systems.ferd.caThe Picture They Paint of YouMusings on the way we frame Coding Assistants, AI SREs, and what this communicates in terms of how these roles are perceived. 0204
Reposted by Fred Hebertjustin from the internet @threlk.net · 21/02/2026this thread is touching on something i find very important: efficiency is sometimes opposed to (or in tension with) other goals efficiency can preclude generalist systems that create flexibility, decentralization and surplus that add resiliency, or non-expert participation that gets people involved 35610
Reposted by Fred HebertNiall Murphy @niallm.bsky.social · 19/02/2026From a discussion in RISF based on the old IBM adage, an updated version for the modern era: 15419
Reposted by Fred HebertLorin Hochstein @norootcause.surfingcomplexity.com · 16/02/2026Whenever the topic of OKRs comes up, I think about Drucker vs Deming. Not a particularly topical thing to write about, but I think it's evergreen. surfingcomplexity.blog/2026/02/16/p...surfingcomplexity.blogPoor Deming never stood a chanceThis post is an elaboration of a shorter post I wrote about five years ago. The two management giants of the mid-twentieth century were Peter Drucker and W. Edwards Deming. Ironically, while Drucke… 0279
Fred Hebert @ferd.ca · 16/02/2026buddy, my challenge isn't generating more content, it's figuring out how to produce a lot less. 0163
Reposted by Fred HebertCat Hicks @grimalkina.bsky.social · 16/02/2026This is my "desert island," most distilled, most succinct piece of advice right now for all the technical people I am talking to who are worried about their learning. 0309
Reposted by Fred HebertCat Hicks @grimalkina.bsky.social · 08/02/2026For the past four years I have seen people say "the decision about what code to write is more important to the code" but does anyone actually look at like, research around what promotes strategic and efficient group decision making? seems like no 7385
Fred Hebert @ferd.ca · 07/02/2026Reading a text on Cognitive Systems Engineering (CSE) & its morality. It starts with the Ea Nasir tablet but then starts going hard and just won't let up, aiming for a morally relativist and nihilistic conclusion of "the real issue is our fake ass sense of absolute morality" The hell is this ride? 171