Fred Hebert @ferd.ca · 27/09/2026Yeah, his concept of just culture is more about drawing a line in the sand where some behavior is clearly unacceptable and punishable, to create a safe area for well-intended workers to operate in and feel comfortable reporting safety issues. It is retributive, and has been debated as well. 030
Fred Hebert @ferd.ca · 27/09/2026In 1997, James Reason published 'Managing Organizational Accidents'. It was influential in popularizing his concepts around the Swiss Cheese model, his specific brand/recipe of safety (and just) culture, etc. and discussed ever since. But it also has this line in it just throwing incredible shade: 23111
Fred Hebert @ferd.ca · 26/09/2026This is a good example of the substitution myth in automation, applied to software workflows. Coactive work and the integration of various functions is different from the individual functions being doable in isolation. 0125
Fred Hebert @ferd.ca · 20/09/2026mornings are starting to be colder which means it is once again time to hover my hands above the toaster while it cooks toast like it’s a campfire 1271
Reposted by Fred HebertLorin Hochstein @norootcause.surfingcomplexity.com · 19/09/2026Today's blog post: a play in one act. surfingcomplexity.blog/2026/09/19/h...surfingcomplexity.blogHow the incident happened: a play in one actA hallway in a tech company office. A and B are standing in the hallway. B is holding a multiple page document. A: Wow, that incident from a couple of weeks ago was a real doozy. Do you know how it… 1114
Fred Hebert @ferd.ca · 19/09/2026There’s a variation you might like based on the law of stretched systems (every system is stretched to operate at its capacity; any improvement will be exploited to achieve a new intensity/tempo of activity) that I believe applies to our ability to understand our systems: ferd.ca/the-law-of-s...ferd.caThe Law of Stretched [Cognitive] SystemsThe law of stretched systems, and how it may also apply to cognitive work, and our ability to deal complexity, such that any improvement is instantly exploited and we forever operate at the edge of un... 070
Fred Hebert @ferd.ca · 19/09/2026but all complex systems are open and all analyses are limited, so it's normal to have to choose boundaries to your analysis, and that selection itself is going to be fairly influential to what is out of scope, what is less controlled, what happens and what emerges, etc. It's a necessary tradeoff. 020
Fred Hebert @ferd.ca · 19/09/2026a key challenge of system modelling is that once your model influences your system over time, you ideally model how it will influence the system and woops, that's a bit harder now, you've added more feedback loops, etc. 110
Fred Hebert @ferd.ca · 19/09/2026tl:dr; yes absolutely; there's decades of literature in other disciplines about encountering these limits (which tech loves to disregard because we think we're built different [and we are to some extent, but in far more restricted ways than we think]), and we tried to use that to inform our methods. 160
Fred Hebert @ferd.ca · 19/09/2026we also took resilience engineering-inspired measures of trying to maintain pace and tempo of these experiments (with real incidents) to build a good improvisational muscle, and sometimes considered basic low-impact incidents as an acceptable tradeoff to avoid worse complexity from a solution. 150
Fred Hebert @ferd.ca · 19/09/2026the chaos engineering view we encouraged was based on exploratory interventions in the system ("what happens if we fuck with x y or z? make your bets then let's see") to pull through the complexity we took for granted, and figure out if our understanding held. But that approach wasn't the only one. 140
Fred Hebert @ferd.ca · 19/09/2026If I can oversimplify, there's been a few branches, including "more flexible control mechanisms let you stay in control and recover" and "no you can't actually plan for everything" which also forks into "we shouldn't build these things" and "actually, adaptive response can help cope even with that." 130
Fred Hebert @ferd.ca · 19/09/2026Increasingly the observation was that systems could fail out of subsystems working as designed, acting reliably and rationally, and still ending in disaster. This lined up with complex systems theories of many kinds, and was a major tension in the 80s going forward. 141
Fred Hebert @ferd.ca · 19/09/2026starting in the late 70s there were rumblings of complexity structurally being its own challenge (Perrow's Normal Accidents is a classic here), which many disciplines have debated to keep things functional and safe. But complexity also kept increasing through tech and globalization, for example. 150
Fred Hebert @ferd.ca · 19/09/2026it taps into a longer tradition of safety that borrows from cybernetics (with principles like "only complexity can cancel complexity") that more or less states that your system gains robustness by accruing complexity. In turn, more complexity gives more room for emergence and undesigned behaviours. 170
Fred Hebert @ferd.ca · 19/09/2026There's a good opinion piece from Amalberti titled "The paradoxes of almost totally safe transportation systems" proposing that what brings you from regular error reduction to very high safety has diminishing returns, and requires control mechanisms making systems inflexible and fragile. 170
Fred Hebert @ferd.ca · 15/09/2026Blame also becomes a problem when it turns into a focal point that distracts from looking at broader contributors, and superficial forms of blamelessness also fall into this trap. I wrote about it more in detail at resilienceinsoftware.org/news/11502437resilienceinsoftware.orgSuperficial Blamelessness | Resilience in Software FoundationIn 2012, John Allspaw (then CTO of Etsy) wrote a seminal blog post on the need for what he called Blameless Postmortems. Built off the notion of a “Just Culture 021
Fred Hebert @ferd.ca · 15/09/2026Blame can play social normative functions to influence people’s behaviour. One element to keep in mind is that actuating blame into shame and punishment is just one way to react. You could also ask the person to be involved in understanding and supporting repairs and improvements to be accountable. 120
Fred Hebert @ferd.ca · 14/09/2026Based on this post's premises, I don't know why there is a need for anyone to have product sense, rather than just letting a good-enough decomposition of the same task analysis be applied to a role that just hasn't gone under the microscope yet. 000
Fred Hebert @ferd.ca · 11/09/2026Every time an e/acc dude tells you about priors, you are expected to interpret this as assumptions based on a statistical context, but it is much more fun to treat it as a legal context where they can’t stop talking about criminal charges they were convicted for. 130
Fred Hebert @ferd.ca · 06/09/2026I’m gardening (though it was incredibly rainy and low touch this year) and doing a masters program on human factors and systems safety part time (which eats ~20h a week easy). 140
Fred Hebert @ferd.ca · 06/09/2026Yeah I did them on and off for 3 non-consecutive years and I think last time I gave up due to one of these “you gotta know the magic theorem ahead of time or this is impossible” problems (that were fairly common in project Euler back then) which felt like a chore and then never came back. 010
Fred Hebert @ferd.ca · 06/09/2026I think it’s now been more than a year since I last wrote software just for fun (with or without AI). It feels more like chores or work by now. It feels like it had been a long time coming, but it just became clearer recently. 2180
Fred Hebert @ferd.ca · 03/09/2026you could describe it as that, or a consequence of the ETTO principle, or just a feature of the information environment... I don't know if the term "error-inducing" is ideal; you could also call it "unforgiving/correction-adverse information environment", which would reframe it as a design issue! 020
Fred Hebert @ferd.ca · 03/09/2026Even this is of limited value since depending on workflow dynamic, people tend to be convinced or swayed by the explanation in front of them. See arxiv.org/abs/2302.12389 and www.nature.com/articles/s41...arxiv.orgExplainable AI is Dead, Long Live Explainable AI! Hypothesis-driven decision supportIn this paper, we argue for a paradigm shift from the current model of explainable artificial intelligence (XAI), which may be counter-productive to better human decision making. In early decision sup... 150
Fred Hebert @ferd.ca · 02/09/2026I'm curious what makes you feel that these principles wouldn't apply to an AI-infused (or AI-centric) system? 040
Reposted by Fred Hebertresilienceinsoftware.org @resilienceinsoftware.org · 01/09/2026Another very cool post you might have missed from our community, on how expertise copes with overload, highlighted by treating incidents as first-class work: resilienceinsoftware.org/news/11453533resilienceinsoftware.orgExpertise and Overload | Resilience in Software FoundationResilience engineering views incidents through a different frame than the conventional approach in the software industry, which tends to treat incidents as an i 022
Fred Hebert @ferd.ca · 01/09/2026If I were a scientist my email signature would absolutely contain the words “more research is needed” 0121
Fred Hebert @ferd.ca · 28/08/2026I personally enjoyed the shape of the work more before the deskilling nature of acceleration kicked in industry-wide. I lean on the part of me that studies joint cognitive systems and resilience engineering because it’s all fascinating to see develop from that point of view. 030
Fred Hebert @ferd.ca · 28/08/2026basically there was no way I'd trust ai to review everything the org did to report relevant data so I did a bunch of deterministic scripts to surface what's likely relevant and made it all diff-friendly so that it was easily reviewable, and once we knew it was regular enough it got bots added. 180
Fred Hebert @ferd.ca · 28/08/2026there can be a ticket for discovery, but then the plan itself has phases prepped that each get ticketed. As new information is gained, the plan is revised and tickets kept in sync. 110
Fred Hebert @ferd.ca · 28/08/2026As we got good at identifying types of changes, it became easier to take the plans we had written for one task and go "this service should be possible to port over by using a plan like <file>" and eventually drift-report auto-tickets migrations we check over. 110
Fred Hebert @ferd.ca · 28/08/2026we ended up writing tools that generate changelog of all APIs or configs we consume (incl. helm values, tf variables, and chart movement across envs), then taking our own repo and shaping it to be easy to compare, to generate "drift reports" informing what to port and check instead of tracking PRs. 110
Fred Hebert @ferd.ca · 28/08/2026well if your plan requires a series of 8 PRs, now you got 9 reviews instead of 8. It's a bit of a preliminary alignment step. 100
Fred Hebert @ferd.ca · 28/08/2026Correction: I don’t know if it’s been reviewed, but it’s been treated as such by some industry folks. 120
Fred Hebert @ferd.ca · 28/08/2026There are literally papers written in reviewed journals about that no longer mattering (for some overly narrow definition of code review). See www.adaptivecapacitylabs.com/2026/08/24/t... for a reference to (and a takedown of) such a paper. What’s intuitive or not may be different to where you sit.adaptivecapacitylabs.comThere is more to code review than (automatable) detection 260
Fred Hebert @ferd.ca · 28/08/2026Yep as with the previous post, and many of my complaining in the recent past, what we knew about systems is still true, it just needs to be applied properly while integrating what’s now available (or in some cases mandated) for leverage. 030
Fred Hebert @ferd.ca · 28/08/2026Right, so formalizing that and leaning into it harder as a mechanism for review-as-model-correction. It’s absolutely not revolutionary, it’s just counterintuitive when in an optimization-centric mindset focuses on code production and bottleneck elimination. 140
Fred Hebert @ferd.ca · 27/08/2026Part of it is context rebuilding. If an engineer plans alone and their PRs are dispatched to 3 coworkers who each reviews a bit, they all have to gradually reconstruct the sequence and plans and you serialize that. If you build awareness and agree before dispatching, ownership+reviewing improves. 160
Fred Hebert @ferd.ca · 27/08/2026I just realized I hadn’t shared this one yet, but I wrote about how the team I’m on restructured its workflows to lean into the code review bottleneck rather than trying to eliminate it when code generation took over, and what that ended up doing: www.honeycomb.io/blog/embraci...honeycomb.ioHow I Came to Embrace the Code Review BottleneckFaced with an endless stream of AI-generated code reviews, our team made the counterintuitive choice to lean into the bottleneck rather than reduce it. 3307
Reposted by Fred HebertLorin Hochstein @norootcause.surfingcomplexity.com · 26/08/2026Every risk register should have an entry that reads “we have misjudged the risks” which is in the “likelihood: high” and “impact: high” region of the risk matrix. 5378
Fred Hebert @ferd.ca · 25/08/2026the public: datacenters’ resource usage is a plague and must be stopped you: there could be more nuance me, a genius manchild: AI datacenters revive Asimov’s sci-fi vision of ever-larger computers encompassing entire cities; it is the pursuit of Multivac and therefore good 120
Fred Hebert @ferd.ca · 20/08/2026Fix code review bottlenecks by doing like private torrent trackers and only allowing people to get their PR reviewed if they reviewed enough PRs beforehand to keep their ratio high enough. 1578
Fred Hebert @ferd.ca · 19/08/2026The first rule of the papers-reading club is that you’re unfortunately already in the papers-reading club and I will send you links and summaries 0140
Fred Hebert @ferd.ca · 17/08/2026Right. I don’t know people were ever disappointed about typing less/in a new box; it’s the nature of what’s typed that connects to skilled work. But the work changes, so does the deskilling/devaluing (IMO), and processing this isn’t simple, particularly when work is part of people’s identity. 020
Fred Hebert @ferd.ca · 17/08/2026Yet again hearing how some skilled work never truly mattered now that it's getting automated, while it absolutely did and was a point of professional pride for many. It's a significant aspect of 'deskilling', which is a known consequence of automation. Pretending otherwise is needlessly unkind. 1357
Fred Hebert @ferd.ca · 17/08/2026this seems more accurate given the frequency of errors right now 0121
Fred Hebert @ferd.ca · 14/08/2026This is the "abstraction-decomposition space". It maps a troubleshooting workflow from a system symptom ("no communication to tape or disk") via a diagonal walk to isolate a faulty component ("short-circuited transistor"). This is such an incredible visualization of so many complex elements! 0124
Fred Hebert @ferd.ca · 14/08/2026The cool thing though is that he then maps a problem-solving path as a line walking that conceptual path within a "conceptual reasoning space" of operators that is event-independent. Neelam Naikar (doi.org/10.1016/j.ap...) re-renders diagrams Rasmussen had made in a few publications in the mid 80s: 140