Root Cause Analysis: A Trigger Event Is Not a Root Cause

Root Cause Analysis: A Trigger Event Is Not a Root Cause

👁4views

A trigger is the ordinary event that set a failure in motion, such as an upgrade, a traffic spike or a wrong value. A root cause is the weakness that allowed that event to become a customer incident. A good RCA must reach that vulnerability, otherwise the conditions for recurrence remain.

Article Summary
  • 1.
    What it is
    Root cause analysis must explain what allowed an event to become a failure, not just what set it in motion. The article separates triggers from latent conditions and shows how to investigate the difference.
  • 2.
    Why it matters
    Naming the trigger lets an organisation close an incident in good faith while leaving every condition for recurrence in place. Reaching the underlying weakness is what makes the fix real.
  • 3.
    Key takeaway
    The most useful RCA question is why this event became a customer incident, not what happened just before the outage.
~14 min read
Listen to this article0 plays

A root cause analysis exists to tell an organisation what it needs to fix, and yet a surprising number of the RCAs I have read over the years tell us something far less useful, which is what happened in the few minutes before a service fell over. An upgrade went out, traffic climbed, a dependency disappeared or somebody typed the wrong value into a console, and that event quietly becomes the explanation for the whole incident. The report gets written, the ticket gets closed and everybody goes back to their day jobs feeling that the matter has been dealt with.

The trouble is that naming the event does not identify the weakness. It explains why the incident happened at that particular moment without explaining why the system was vulnerable to it, and unless the analysis reaches that vulnerability, the organisation can close the incident in good faith while leaving every condition for its recurrence exactly where it found them. The trigger explains what set the failure in motion; the root cause analysis must explain what allowed it to become a failure. Sometimes a change does more than expose an old weakness and actually introduces a new one, such as a defective release or a configuration that was simply wrong, and in that case the defect belongs in the findings, but the analysis still has to explain why it reached production and why its effects were not contained.

1. Why the trigger cannot be the whole answer

Triggers come in every shape, from the entirely routine, such as demand rising through the day or a process restarting, to the genuinely exceptional, such as a regional power failure or a malformed message nobody had imagined. How unusual the trigger was does not change the job of the analysis, which is to identify the conditions that allowed it to cause harm and to say whether they fell inside or outside the limits the service was supposed to handle. None of this is new: James Reason distinguished active failures at the sharp end from latent conditions created by design and organisational decisions, which lie dormant until something exposes them (Reason, 2000), and Richard Cook observed that overt failure in complex systems generally requires multiple faults, each insufficient on its own (Cook, 1998). You do not need to abandon the phrase “root cause analysis” to take that seriously, but you do need to stop treating the most recent event in the timeline as the answer.

2. Two outages and two incomplete answers

Consider an outage during a platform upgrade. Recording it as “change related” may be a useful classification, but it is not an explanation, because the RCA still has to tell us why the service could not tolerate a planned change. Perhaps the deployment lacked the redundancy to lose a node gracefully, perhaps the application dropped in flight requests when its instances were drained, or perhaps the upgrade itself introduced a defect; those are hypotheses to investigate, and the RCA should report whichever of them the evidence supports.

Now consider an incident attributed to “larger than expected load” and recovered by somebody manually changing a fixed number of replicas. The service was running at a fixed capacity that could not respond to demand on its own, and the only way to adapt it was for a person to notice and change a number by hand. That is not an inference drawn from the recovery but a fact about how the service was configured, and it is the weakness the RCA needs to name. A fixed replica count is really a guess about future demand written into configuration, and guessing load is a bad business to be in. The whole point of running on an elastic platform such as Kubernetes on EKS is that you should not have to work out in advance how much capacity you need, because the platform scales with demand within limits you have chosen deliberately, so if someone has to adjust replica counts by hand because the load was “unexpected”, something is wrong with the design rather than the forecast. The extra demand was simply the trigger, and a marketing campaign, a month end peak or a retry storm would have found the same weakness just as easily. The investigation should still check whether anything else contributed, such as a release that increased the resources each request consumes, but it should certainly not stop at the observation that more customers arrived, since for most businesses that is the entire point.

3. Timelines and labels are not findings

Both examples point to a broader problem, which is that organisations easily mistake an incident timeline for an explanation. A timeline establishes the sequence of events, but an RCA has to identify the conditions, decisions and missing or ineffective controls that allowed those events to harm customers, and the most useful question I know for getting there is to ask why this event became a customer incident. If a server restarted, why did the service depend on that server staying up? If someone entered the wrong value, why did the process allow it to reach production, and why were the consequences not contained to something smaller than the whole service?

Familiar labels deserve the same scepticism. “Configuration issue”, “human error” and “insufficient testing” are categories that could describe thousands of unrelated failures, whereas a finding names which configuration was wrong, which safeguard was missing and which behaviour the testing failed to exercise. This is also why I am wary of “five whys” applied mechanically; Alan Card argues that it tends to follow a single causal path, stops at an arbitrary depth and can be shaped by what the people running it already believe (Card, 2017). Asking why repeatedly is a perfectly good prompt, but real failures branch, so the questioning should follow every branch the evidence opens and stop at a weakness that can be fixed rather than at whatever the fifth answer happens to be. None of this means every conceivable event must be survivable, because systems have finite capacity, budgets and agreed service levels, but it does mean the investigation must say where the design, implementation or operating controls fell short, including whether the assumed limits were sensible in the first place. Occasionally the honest finding is that nobody had ever decided what the limit was.

4. Contributing factors are not the root cause either

Between the trigger and the weakness sits a third category that causes almost as much confusion. Contributing factors are the things that made the outage more likely, worse or longer without being the reason the system was vulnerable. In the fixed replica example, perhaps the latency alert fired twenty minutes late, the on call engineer lost time finding where the replica count was configured, the runbook was out of date and escalation to the owning team was slow. Each of those deserves an action, but if all of them were fixed tomorrow the service would still run at a fixed capacity that cannot respond to demand, and the next spike would still cause an outage; the organisation would simply detect it sooner and recover faster.

The test I find most useful is to ask, for each factor, whether fixing it on its own would stop the same trigger causing the same failure. If the outage would still happen, only shorter or smaller, it is a contributing factor; if fixing it would remove or contain the whole class of failure, you are close to the weakness the RCA needs to name. The distinction matters because contributing factors are usually far easier to fix, and tuning an alert or rewriting a runbook can be closed off in an afternoon while redesigning how a service scales may take weeks and a difficult conversation about priorities. An action list made up entirely of detection and response improvements can look thorough while doing nothing about why the service failed. The line is not always sharp, and where several conditions are each necessary for the outage, as Cook would expect, the honest answer is to name them together as the weakness rather than let a long list of small findings disguise the absence of the important one.

Amazon’s Correction of Error process, the COE, makes this separation easy to see. Its incident questions are grouped into detection, diagnosis, mitigation and prevention, alongside an impact assessment, a timeline and action items with owners and dates, and it is explicitly not a process for finding someone to blame (Caro, Ossa and Hanley, 2022). Detection, diagnosis and mitigation are where contributing factors naturally live; prevention is where the weakness has to appear. A COE whose prevention answers amount to “forecast better” or “freeze changes” has documented the trigger and the contributing factors and stopped short of the cause, however complete the rest of the document looks.

5. What looking past the trigger looks like in practice

The summary AWS published after the February 2017 S3 disruption in US-EAST-1 is a good example. According to that summary, an authorised team member following an established playbook entered one input to a command incorrectly and removed far more servers than intended, including servers supporting subsystems S3 depends on. AWS could have stopped at operator error, but its remediation was aimed at the tool: it was changed to remove capacity more slowly and to refuse to take a subsystem below its minimum required capacity, and other operational tools were audited for similar checks (AWS, 2017). The mistyped input was the trigger; the weakness was a tool that let a single input remove too much capacity too quickly.

CrowdStrike’s analysis of the July 2024 Channel File 291 incident shows the other case, where a change introduced the defect. Beneath the headline of a faulty content update sat findings an engineering team could act on: a template type defined 21 input fields while the sensor code supplied only 20, the mismatch was not caught by validation or testing, and the content interpreter performed an out of bounds memory read when it reached for the twenty first input. The remediation included checks on the number of inputs and staged deployment of this content rather than release to every sensor at once (CrowdStrike, 2024). Each of those findings can be fixed and tested, which is exactly what “bad update” cannot.

6. The real cost of chasing triggers

The wording of an RCA decides where effort goes, and when a report names a trigger, the natural instinct is to control the trigger. If an upgrade was to blame, the organisation introduces a change freeze, adds an approval step and quietly does fewer upgrades; if unexpected demand was to blame, it invests in better forecasting, which is really just a more elaborate way of guessing load. These responses are visible, quick and look responsible in a board pack, but they carry real risks.

The first is that the weakness stays exactly where it was. A service that cannot survive a node being drained during an upgrade cannot survive a hardware failure, a provider maintenance event or a kernel panic either, and nobody can freeze those, so the freeze removes only the triggers the organisation controls and postpones the outage rather than preventing it. The second is that deferred upgrades accumulate: each skipped upgrade makes the next one larger and riskier, while the platform falls behind on security patches and drifts towards versions out of vendor support, trading a visible operational risk for a less visible security risk that is often worse. The third is batching, because work that queues behind a freeze ships in larger bundles with a bigger blast radius and is harder to diagnose when it fails. DORA’s 2019 research found no evidence that formal external change review was associated with lower change failure rates, and found heavyweight approval associated with worse delivery performance, recommending peer review supported by automated checks instead (DORA, 2019; see also Forsgren, Humble and Kim, 2018). That is correlational survey evidence rather than proof that change boards cause instability, but it should make us cautious about treating more approval as the obvious fix.

The fourth risk is that the organisation slowly loses the ability to change safely at all, because pipelines, rollback paths and runbooks decay when they are rarely exercised, and the moment that is discovered is usually an urgent patch for a critical vulnerability. The fifth is cultural: when every incident produces more paperwork and less change, engineers learn to classify incidents in ways that avoid the freeze rather than explain the failure, while the organisation closes its action items and reports progress that has not happened. A short freeze can still be sensible containment while a known weakness is fixed, provided it has an exit criterion tied to that fix, such as lifting it once graceful shutdown has been proven under load. The goal is to make change safe rather than rare, through redundancy, graceful shutdown, canary releases, staged rollouts and automated rollback, until a routine upgrade becomes boring again.

7. Recovery is not prevention, and the fix must follow the weakness

Recovery creates the same illusion. Rolling back, restarting or adding capacity is often exactly right during an incident, but none of it necessarily removes the vulnerability, so a completed recovery action should never be treated as a completed preventive action; if most of an RCA’s actions were already done on the incident call, it has documented the firefight rather than the cause. The corrective action must follow from the weakness the evidence actually established and come with proof that it works. If shutdown dropped active requests, fix shutdown and demonstrate requests surviving a drain under realistic load. If a fixed replica count could not follow demand, replace it with automatic scaling, with sensible minimums, maximums and a signal that reflects real load, and show it coping with a replay of the demand that broke it. If a tool let one input remove too much capacity, put an enforced limit in the tool rather than a reminder in the runbook.

8. Blameless does not mean vague

None of this conflicts with a blameless culture. The Google SRE book describes a blameless postmortem as one that identifies contributing causes without indicting any individual or team, assuming people acted with good intentions on the information they had (Lunney and Lueder, 2016). That is not a licence for vague findings; we can examine decisions, ownership and control failures precisely without passing judgement on anyone’s character. If anything, precise findings are what make accountability practical, because they say what must change, who will change it and how success will be shown, whereas nobody can be held responsible for fixing “a configuration issue”.

9. Before you sign off an RCA

Before accepting a root cause analysis, I would work through these questions and be honest about the answers:

  1. Does the conclusion name a weakness the organisation can correct, or only an event?
  2. Does it explain why that event became a customer incident?
  3. Are contributing factors separated from the weakness, and is at least one action aimed at the weakness itself rather than at detecting or recovering from it faster?
  4. Does each action remove or contain the weakness, rather than restate the recovery?
  5. Is there evidence, such as a test, a replay or a game day, that the action works?
  6. If the actions include a freeze or extra approvals, is there an exit criterion tied to a specific fix?
  7. Would the same weakness still be exposed by a different trigger?

The trigger and the weakness both belong in the causal account, because one explains when and how the failure started and the other explains why it was able to cause harm. What matters is that the analysis does not stop at the first of them, and carries on until it can say what was weak and what we are doing to make it stronger.

References

  1. Reason, J. (2000). Human error: models and management. BMJ, 320(7237), 768-770.
  2. Cook, R. I. (1998, revised 2000). How Complex Systems Fail. Cognitive Technologies Laboratory, University of Chicago.
  3. Card, A. J. (2017). The problem with ‘5 whys’. BMJ Quality & Safety, 26(8), 671-677.
  4. Caro, L., Ossa, J. and Hanley, J. (2022). Why you should develop a correction of error (COE). AWS Cloud Operations Blog.
  5. Amazon Web Services (2017). Summary of the Amazon S3 Service Disruption in the Northern Virginia (US-EAST-1) Region.
  6. CrowdStrike (2024). Channel File 291 Incident: Root Cause Analysis is Available.
  7. DORA (2019). Capabilities: Streamlining change approval, drawing on the 2019 Accelerate State of DevOps Report.
  8. Forsgren, N., Humble, J. and Kim, G. (2018). Accelerate: The Science of Lean Software and DevOps. IT Revolution.
  9. Lunney, J. and Lueder, S. (2016). Postmortem Culture: Learning from Failure. In Site Reliability Engineering, O’Reilly.

Leave a comment

Your email address will not be published. Your first comment is held for approval.