Sign in

Fred Hebert

@ferd.ca
2.1K followers 190 following 560 posts

Principal SRE @ honeycomb.io, Tech Book Author, Resilience in Software Foundation board member, Erlang Ecosystem Foundation co-founder, Resilience Engineering fan. SRE-not-sorry. blog: ferd.ca notes: ferd.ca/notes

PostsRepliesMedia
Fred Hebert @ferd.ca · 27/09/2026
In 1997, James Reason published 'Managing Organizational Accidents'. It was influential in popularizing his concepts around the Swiss Cheese model, his specific brand/recipe of safety (and just) culture, etc. and discussed ever since. But it also has this line in it just throwing incredible shade:
Text from a book:

For example, some individuals are considerably more absentminded than others. If the person in question has a previous history of unsafe acts, it does not necessarily bear upon the culpability of the error committed on this particular occasion, but it does indicate the necessity for corrective training or even career counselling along the lines of 'Don't you think you would be doing everyone a favour if you considered taking on some other job within the company?'. This is the way that management acquires some of its most distinguished members. Absentmindedness has nothing whatsoever to do with ability or intelligence, but it is not a particularly helpful trait in a pilot or control room operator.
[...]

"This is the way management acquires some of its most distinguished members" is manually highlighted and a '!!' is in the margin.
23011
Fred Hebert @ferd.ca · 26/09/2026
This is a good example of the substitution myth in automation, applied to software workflows. Coactive work and the integration of various functions is different from the individual functions being doable in isolation.
0125
Fred Hebert @ferd.ca · 20/09/2026
mornings are starting to be colder which means it is once again time to hover my hands above the toaster while it cooks toast like it’s a campfire
1271
Reposted by Fred Hebert
Lorin Hochstein @norootcause.surfingcomplexity.com · 19/09/2026
Today's blog post: a play in one act. surfingcomplexity.blog/2026/09/19/h...
surfingcomplexity.blog
How the incident happened: a play in one act
A hallway in a tech company office. A and B are standing in the hallway. B is holding a multiple page document. A: Wow, that incident from a couple of weeks ago was a real doozy. Do you know how it…
1114
Fred Hebert @ferd.ca · 06/09/2026
I think it’s now been more than a year since I last wrote software just for fun (with or without AI). It feels more like chores or work by now. It feels like it had been a long time coming, but it just became clearer recently.
2180
Reposted by Fred Hebert
resilienceinsoftware.org @resilienceinsoftware.org · 01/09/2026
Another very cool post you might have missed from our community, on how expertise copes with overload, highlighted by treating incidents as first-class work: resilienceinsoftware.org/news/11453533
resilienceinsoftware.org
Expertise and Overload | Resilience in Software Foundation
Resilience engineering views incidents through a different frame than the conventional approach in the software industry, which tends to treat incidents as an i
022
Fred Hebert @ferd.ca · 01/09/2026
If I were a scientist my email signature would absolutely contain the words “more research is needed”
0121
Fred Hebert @ferd.ca · 27/08/2026
I just realized I hadn’t shared this one yet, but I wrote about how the team I’m on restructured its workflows to lean into the code review bottleneck rather than trying to eliminate it when code generation took over, and what that ended up doing: www.honeycomb.io/blog/embraci...
honeycomb.io
How I Came to Embrace the Code Review Bottleneck
Faced with an endless stream of AI-generated code reviews, our team made the counterintuitive choice to lean into the bottleneck rather than reduce it.
3307
Reposted by Fred Hebert
Lorin Hochstein @norootcause.surfingcomplexity.com · 26/08/2026
Every risk register should have an entry that reads “we have misjudged the risks” which is in the “likelihood: high” and “impact: high” region of the risk matrix.
5378
Fred Hebert @ferd.ca · 20/08/2026
Fix code review bottlenecks by doing like private torrent trackers and only allowing people to get their PR reviewed if they reviewed enough PRs beforehand to keep their ratio high enough.
1578
Fred Hebert @ferd.ca · 19/08/2026
The first rule of the papers-reading club is that you’re unfortunately already in the papers-reading club and I will send you links and summaries
0140
Fred Hebert @ferd.ca · 17/08/2026
Yet again hearing how some skilled work never truly mattered now that it's getting automated, while it absolutely did and was a point of professional pride for many. It's a significant aspect of 'deskilling', which is a known consequence of automation. Pretending otherwise is needlessly unkind.
1357
Fred Hebert @ferd.ca · 17/08/2026
this seems more accurate given the frequency of errors right now
the github unicorn from its error pages, without the horn.
0121
Fred Hebert @ferd.ca · 14/08/2026
Although there have been changes in cognition and systems theory, I wanted to bring up some stuff Rasmussen published in the 80s that was really elegant. I'll use 3 diagrams he published, covering the ideas of abstraction hierarchies and how people operate systems when troubleshooting them.
1136
Fred Hebert @ferd.ca · 10/08/2026
This is a very good text on comparing ecology and organizational dynamics when it comes to harvesting signals; big fan of this one. psychsafety.com/a-practical-...
psychsafety.com
Organisational Indicator Species: being an Organisational Ecologist
The more elaborate and expensive our system for understanding something, the less we may actually understand it. Ecologists face the same problem — you can't measure the health of an ecosystem directl...
073
Fred Hebert @ferd.ca · 10/08/2026
Wrote up a bunch of stuff about some patterns in system design, about the tension, contrast, and possibility of composing approaches of analytical decomposition to increase control, and of complexity-aware stances for emergence, and some pitfalls of either stance: ferd.ca/control-and-...
ferd.ca
Control and complexity: tension in systems design
composing two broad approaches, one based on analytical decomposition that aims to maintain control over a system, and one based on a perspective of complex systems that resist analysis, and implicati...
0177
Reposted by Fred Hebert
Lorin Hochstein @norootcause.surfingcomplexity.com · 03/08/2026
Quick little brainstorm-y blog post about traditional versus resilience engineering focuses (foci?): surfingcomplexity.blog/2026/08/02/t...
surfingcomplexity.blog
Traditional versus resilience engineering views
As a fan of resilience engineering, I often differ with people on where we should focus our scarce engineering cycles in order to improve reliability. I thought it would be a useful exercise to bra…
172
Fred Hebert @ferd.ca · 24/07/2026
Delegating pressure to the final individual, who is now tasked with continuously fixing the entire system’s misalignments through their personal choices.
2254
Fred Hebert @ferd.ca · 14/07/2026
I was a reviewer for this book so I can give you some spoilers about it: it rules, get a copy of it, start telling everyone about cumulative culture, unbuild the myths that keep your teams down, make shit be meaningful.
1288
Fred Hebert @ferd.ca · 13/07/2026
One of the ironies about AI agents in ops tasks is that it feels like there has never been as much interest in creating a forgiving environment with proper structural support than through promising to remove humans from it, finally forcing a less individualistic and blameful approach to design.
7478
Fred Hebert @ferd.ca · 10/07/2026
at this point why not just host my repos on the staging servers, github
a search box from the github UI showing the word 'flag' typed in, and then a tooltip displaying '*flag* will be ignored since log searching is not yet available'
081
Fred Hebert @ferd.ca · 03/07/2026
These Angine de Poitrine fans got the best flags I’ve seen in a while; that’s some amazing concert gear.
Angine de poitrine at La Noce 2026 on stage, with two fans in the crowd waving large Quebec flags where the blue squares are replaced by black with white polka dots, the fleur de lys are golden, and the white is covered with black polka dots to match the band’s aesthetic.
0262
Fred Hebert @ferd.ca · 23/05/2026
Even after years, one of the weirdest parts of gardening to me is needing to harden the seedlings before transplanting them. Like “yes hold on a minute I gotta take the plants out so they can play outdoors for a while, but they gotta be in before streetlights turn on” is a real and necessary thing.
0100
Fred Hebert @ferd.ca · 19/05/2026
The ongoing stream of software engineering pieces that mention that the future is in writing spec but never bother to define what a specification is or at what abstraction levels it should be is appalling; arguably, tickets are a spec, the code is a spec, and work between both is connecting dots.
3315
Fred Hebert @ferd.ca · 17/05/2026
“decisions with lasting social consequences are attached to a future moment when the technology is assumed to have revealed its true form, rather than addressed in the present in which it already operates. […] it defers responsibility by attaching accountability to a moment that never materializes”
link.springer.com
Waiting for AGI - AI & SOCIETY
AI & SOCIETY -
03210
Reposted by Fred Hebert
Lorin Hochstein @norootcause.surfingcomplexity.com · 10/05/2026
New blog post about flipping the bozo bit: surfingcomplexity.blog/2026/05/09/f...
surfingcomplexity.blog
Flipping the bozo bit on flips the learning off
I’m too young to have seen Bozo the Clown myself, but I’m old enough to get the references “Flipping the bozo bit” is an expression from the software world. Think about a ti…
1114
Fred Hebert @ferd.ca · 02/05/2026
“[…] we reached for Recon, an amazing tool for diagnosing issues […] (the related Erlang in Anger is more-or-less required reading as all on-call engineers end up scouring its pages eventually)” Wild! I wrote these 10 years ago to help coworkers, and they’re still useful for real world issues now!
discord.com
You’ve Got (Too Much) Mail: Behind the Scenes of the 3/25/26 Voice Outage
On March 25th, voice and video on Discord suffered major degradation beginning at 12:13 PDT, lasting a little over three hours. Learn how the issue originated, how it affected systems across Discord, ...
0192
Fred Hebert @ferd.ca · 29/04/2026
Infinite love to my SRE coworkers, one of whom casually dropped this single line in a retro: "the industrial hourly deploy train and its consequences have been a disaster for society"
15717
Fred Hebert @ferd.ca · 20/04/2026
Writing my action items on a piece of paper that I will burn to ward off spirits
092
Fred Hebert @ferd.ca · 18/04/2026
I wrote for @resilienceinsoftware.org on "Superficial Blamelessness", where under the label of "blamelessness", we avoid punishing people, yet still focus fixes and interventions based on the same individualistic framing rather than a broader systemic stance. resilienceinsoftware.org/news/11502437
resilienceinsoftware.org
Superficial Blamelessness
In 2012, John Allspaw (then CTO of Etsy) wrote a seminal blog post on the need for what he called Blameless Postmortems. Built off the notion of a “Just Culture” from the research of Sidney Dekker, he...
0207
Fred Hebert @ferd.ca · 08/04/2026
I've gotten an early copy of Crisis Engineering by Marina Nitze, Matthew Weaver, and Mikey Dickerson, and just in time for its release today, here's my review of it: ferd.ca/notes/on-cri... TL:DR; I like it, good perspectives and an interesting mix of approaches.
092
Fred Hebert @ferd.ca · 07/04/2026
One of my assignments asks whether a given set of psychological constructs amount to “good science” and it’s very weird as someone who’s written zero papers and done no real science to just attempt to go “sure dawg it’s okay science I guess but these hundreds of researchers could pick better models”
140
Fred Hebert @ferd.ca · 03/04/2026
Every time I am faced with this Swedish login form, I have to say 'Logga in' out loud with a terminator voice. It is one of the few rules we can't change and must be respected.
swedish login from with a button that says 'Logga in'
2271
Fred Hebert @ferd.ca · 23/03/2026
hell yeah @resilienceinsoftware.org swag is in! the law of requisite variety states that only variety in the regulator can destroy variety in the system being regulated. so if you need to deal with complexity you know you gotta join the club & begrudgingly increase complexity to keep things simple
Black hoodie showing 'anti complexity complexity club' in a red explosion based on the logo of the Resilience in Software Foundation.
1152
Fred Hebert @ferd.ca · 16/03/2026
would you rather have many incidents of various small to moderate size for the foreseeable future or just one very big incident and then be done with them for good?
880
Fred Hebert @ferd.ca · 08/03/2026
As usual @grimalkina.bsky.social is worth listening to. These principles are also worth considering and applying in all sorts of contexts. Here’s a sample from safety research I happened to read just yesterday (Dekker - Reconstructing human contributions to accidents, 2002) that aligns with it!
What is striking about many accidents in complex systems is that people were doing exactly the sorts of things they would usually be doing the things that usually lead to success and safety. Mishaps are more typically the result of everyday influences on everyday decision making than they are isolated cases of erratic individuals behaving unrepresentatively.
People are doing what makes sense given the situational indications, operational pressures, and organizational norms existing at the time. Accidents are seldom preceded by bizarre behavior. People's errors and mistakes (such as there are in any objective sense) are systematically coupled to their circumstances and tools and tasks. [...] What people do makes sense to them at the time—it has to, otherwise they would not do it. People do not come to work to do a bad job; they are not out to crash cars or airplanes or ground ships. The local rationality principle says that people do things that are reasonable, or rational, based on their limited knowledge, goals, and understanding of the situation and their limited resources at the time. […] failures are baked into the nature of people's work and organization; that they are symptoms of deeper trouble or by-products of systemic brittleness in the way business is done. It means having to find out what people did back there and then actually make sense of it given the organization and operation that surrounded them.
To explain outcome failure, it is necessary to convert the search for human failures into a search for human sensemaking. The question is not "where did people go wrong?" but "why did this assessment or action make sense to them at the time?" Such real insight is derived not from judging people from the position of retrospective outsider, but from seeing the world through the eyes of the protagonists at the time. When looking at the sequence of events from this perspective, a very different story often struggles into view.
2295
Fred Hebert @ferd.ca · 07/03/2026
When I joined the program I'm currently in, I told myself I would not read more papers and technical books outside of it because I'd need to balance about my energy levels—not spending it all on this. I was right, but also there's lots of other cool nerd shit I want to read through now and welp.
0120
Reposted by Fred Hebert
Liz Fong-Jones (方禮真) @lizthegrey.com · 04/03/2026
Here is the fuller writeup I promised, a little bit overdue (I said Jan but it actually went out in Feb). Credit to @ferd.ca for the writeup, as well as to all of our incident responders who worked this 12+ day incident. www.honeycomb.io/blog/inciden...
honeycomb.io
Incident Report: Exercises, Cleanups, and Evacuations
Every year, Honeycomb runs disaster recovery scenarios in multiple environments, including in production. Although each of our instances runs in a single region, on at least three Availability Zones ...
0122
Fred Hebert @ferd.ca · 01/03/2026
I really need to do this year’s garden planning. It’s gonna be time for seedlings in a few weeks and it’s gonna be nice to once again get going on one of them hobbies where you can’t really obsessively dictate the pace nor feel pressured in going faster. Just watch the plants grow and see what goes.
1150
Fred Hebert @ferd.ca · 26/02/2026
Back in December, we had a large outage at work. The internal investigation took a while and the internal report was roughly 40 pages long. For the public, we managed to try and condense it to a much shorter format that we think can still offer useful insights to other organizations:
honeycomb.io
Incident Report: Exercises, Cleanups, and Evacuations
Every year, Honeycomb runs disaster recovery scenarios in multiple environments, including in production. Although each of our instances runs in a single region, on at least three Availability Zones ...
1124
Fred Hebert @ferd.ca · 25/02/2026
I believe I had a relatively intuitive sense of how much Swiss cheese I could consume in one sitting before I can take no more. I also believe I now have a fairly empirical sense of how much literature about the Swiss Cheese Model I can consume in one sitting before I can take no more.
3170
Reposted by Fred Hebert
Chastity Blackwell @blackisis.bsky.social · 23/02/2026
I also think this discussion about how SRE work is being devalued by these products is at least parallel to the discussion about how people are willing to write clear documentation for *AI* consumption, but never put any value on it when it was for *actual people*.
082
Fred Hebert @ferd.ca · 23/02/2026
AI SREs are framed as nameless job automation; Coding Assistants as named partners. The framing used to create these products reveals a lot about how the builders and buyers perceive these roles. I also write about the challenges and risks of picking self-limiting analogies in building systems.
ferd.ca
The Picture They Paint of You
Musings on the way we frame Coding Assistants, AI SREs, and what this communicates in terms of how these roles are perceived.
0204
Reposted by Fred Hebert
justin from the internet @threlk.net · 21/02/2026
this thread is touching on something i find very important: efficiency is sometimes opposed to (or in tension with) other goals efficiency can preclude generalist systems that create flexibility, decentralization and surplus that add resiliency, or non-expert participation that gets people involved
35610
Reposted by Fred Hebert
Niall Murphy @niallm.bsky.social · 19/02/2026
From a discussion in RISF based on the old IBM adage, an updated version for the modern era:
15419
Reposted by Fred Hebert
Lorin Hochstein @norootcause.surfingcomplexity.com · 16/02/2026
Whenever the topic of OKRs comes up, I think about Drucker vs Deming. Not a particularly topical thing to write about, but I think it's evergreen. surfingcomplexity.blog/2026/02/16/p...
surfingcomplexity.blog
Poor Deming never stood a chance
This post is an elaboration of a shorter post I wrote about five years ago. The two management giants of the mid-twentieth century were Peter Drucker and W. Edwards Deming. Ironically, while Drucke…
0279
Fred Hebert @ferd.ca · 16/02/2026
buddy, my challenge isn't generating more content, it's figuring out how to produce a lot less.
0163
Reposted by Fred Hebert
Cat Hicks @grimalkina.bsky.social · 16/02/2026
This is my "desert island," most distilled, most succinct piece of advice right now for all the technical people I am talking to who are worried about their learning.
0309
Reposted by Fred Hebert
Cat Hicks @grimalkina.bsky.social · 08/02/2026
For the past four years I have seen people say "the decision about what code to write is more important to the code" but does anyone actually look at like, research around what promotes strategic and efficient group decision making? seems like no
7385
Fred Hebert @ferd.ca · 07/02/2026
Reading a text on Cognitive Systems Engineering (CSE) & its morality. It starts with the Ea Nasir tablet but then starts going hard and just won't let up, aiming for a morally relativist and nihilistic conclusion of "the real issue is our fake ass sense of absolute morality" The hell is this ride?
Engineering the Morality of Cognitive Systems and CSE; Figure 5.1 showing the Ea Nasir tablet with its full translation:

> Tell Ea-nasir: Nanni sends the following message. "When you came, you said to me as follows: 'I will give Gimil-Sin (when he comes) fine quality copper ingots.' You left then, but you did not do what you promised me. You put ingots that were not good before my messenger (Sit-Sin) and said: 'If you want to take them, take them; if you do not want to take them, go away!' What do you take me for, that you treat somebody like me with such contempt? I have sent as messengers gentlemen like ourselves to collect the bag with my money (deposited with you but you have treated me with contempt by sending them back to me empty-handed several times, and that through enemy territory. Is there anyone among the merchants who trade with Telmun who has treated me in this way? You alone treat my messenger with contempt! On account of that one (trifling) mina of silver which I owe?) you, you feel free to speak in such a way, while I have given to the palace on your behalf 1080 pounds of copper, and umi-abum has likewise given 1080 pounds of copper, apart from what we both have had written on a sealed tablet to be kept in the temple of Samas. How have you treated me for that copper? You have withheld my money bag from me in enemy territory; it is now up to you to restore (my money) to me in full. Take cognizance that (from now on) I will not accept here any copper from you that is not of fine quality. I shall (from now on) select and take the ingots individually in my own yard, and I shall exercise against you my right of rejection because you have treated me with contempt."As computational capacities have increased, the residual functions remaining to the human operator have been largely a litany of those tasks that computers had yet to be engineered so as to be able to accomplish. This left human beings, in Kantowitz's most evocative phrase, as the sub-systems of last resort. The moral concern with respect to a human operator's choice in terms of their own personal expectations, aspirations, dignity, and enjoyment of work was largely neglected, ignored, or deferred to the technical discussion of evolving automation capacities. Human hopes, desires, or more generally the hedonomic dimensions of work have very rarely been emphasized as imperatives, where the evolving technical capacities of the computer system have almost without exception served to drive the debate.Up until the emergence of CSE, many "systems-based" errors were thought of as the proximal problem of the closest human operator who was in some fashion error- prone or in other ways limited, biased, or even worse somehow "deserving" of such a fate (Arbous and Kerrich 1951; Dekker 2007b; Flach and Hoffman 2003; Wilder 1927). These have been couched in psychological terms such as lacks of attention, failures of situation awareness, and so on. It might be argued that these are somewhat more benign labels than one tellingly derived from the English military for pilot error: that is, a lack of moral fiber. These ways of stating the problem directly entail a stance that it is necessarily the advancements in technology that will circumvent such failures (Hancock 2000). Just occasionally, there arise assertions that collective social governances also bear some responsibility for such mishaps.

Blame, the moral pronouncement and disapprobation of the powerful collective on the behavior of the poor or poor individual or indeed any other less powerful collective, remains an intrinsic result of this divide. Here, error (sin, and its correlate-evil) becomes an inherent property or characteristic of that identified individual or minority. The excision of that one deviant operator or small group of operators quickly follows and thus the standard narrative is sustained. This notion echoes the disease model of harm in getting rid of the defect or one individual source. One great contribution of CSE has always been to look to expose this myth and to seek an alternative perspective...
Our western morality here emanates from an essential schizophrenia that derived from the concept of a munificent but omnipotent deity. [...] [Law as Western moralisty] has a vested interest in the status quo and so no great concern for changing it. Thus, CSE may be embraced by science and some limited segments of the business community but has yet to exert any proportionate social and moral impact because of the intransigence of the identified inertial institutions. [...]

5.9 CONCLUSIONS
As a result of the historical antecedents, which in the western world have divided process from purpose, our present circumstances generate complex work systems of great technological sophistication but erected upon ethical bases that are implicit, inconsistent, incomplete, and ultimately inoperable. Utilitarian, patchwork moral principles, derived from a dissolving base of religious dogma, are now morphing into ad hoc, secular legal arcana whose appeal to various poorly defined and developed human delusions, such as the notion of error, free will, and the omnivorous drive for greater efficiency, continues to subserve the dominant segment of profit-based capitalism (and polemically, see Dawkins 2009). [...]

Moral relatively has to be recognized, yet contemporary pragmatism employs and emphasizes an "as if" policy such that the search for moral certitude continues. Like error itself, morality is a convenient illusion that helps sustain societies in their necessary collective delusion. Morality is an intimate part of the narrative that we have always told ourselves (Homer 735 BC). [...]. Morality expresses social approbation and disapprobation and CSE subserves these moral goals (as well as other forms of purpose). What is good or bad, correct or incorrect are collective interpretations (most often after the fact) that support the individual and collective narrative that we have chosen as our conduit to collective reality (and, of course, see Nietzsche 1886).
171