Gross Risk vs Net Risk: A Technology Leader's Guide to Why Freezing Change Doesn't Make Legacy Systems Safer

Gross Risk vs Net Risk: A Technology Leader’s Guide to Why Freezing Change Doesn’t Make Legacy Systems Safer

πŸ‘11views

Freezing change does not make legacy systems safer because it maximises net risk. When a team stops understanding cause and effect, operation turns into superstition and the system resembles a Jenga tower where every change feels dangerous. Real risk management means understanding the environment deeply, then reducing gross risk through controlled changes such as retiring obsolete components and eliminating single points of failure.

CloudScale AI SEO: Article Summary
  • 1.
    What it is
    Gross risk vs net risk explains why technology leaders who freeze all change on fragile legacy systems often choose the most dangerous strategy available. The article teaches how to build a causal model of an environment before managing its risk.
  • 2.
    Why it matters
    It argues that celebrating incident free months while narrating outcomes instead of managing underlying probability destroys stakeholder trust, and that honest bankruptcy declarations plus deep technical understanding earn the right to lead recovery.
  • 3.
    Key takeaway
    A bad outcome that permanently reduces future uncertainty can be more valuable than a good outcome that teaches you nothing.
~19 min read
🎧 Listen to this article

I have been parachuted into a number of troubled technology environments during my career, and they usually shared the same characteristics: too many incidents, nervous business stakeholders, old technology, fragile architecture, and a technology team that had gradually lost confidence in its own ability to change anything.

One of the first things I learned was not to rush into stakeholder meetings and start reassuring people. I wanted to understand the object first. I would spend a disproportionate amount of time going through telemetry, architecture, infrastructure, storage, operating systems, database versions, network paths, dependencies, batch schedules, failure history and incident data, trying to understand the environment almost switch by switch. What is this thing? How does it actually work? Where is it fragile? Where are we operating outside design tolerances? What happens when this component fails, and how would we recover it? Which parts of our architecture do we truly understand, and which parts are simply still running because nobody has disturbed them recently?

There was a reason for this obsession. You cannot manage the risk of something you do not understand. And underneath everything that follows sits one idea: the risk of doing something cannot be evaluated independently of the risk of continuing to do nothing.

1. Start By Declaring Yourself Bankrupt and Clueless

Before any of that technical work begins, there is something more basic you have to get right, and it has nothing to do with architecture. Be honest.

When you walk into a struggling environment, the temptation is to project confidence. You want to stand in front of nervous stakeholders and tell them it’s under control, that you’ve seen this before, that you know exactly what to do. Resist that. Don’t be arrogant about it, and don’t pretend to a certainty you don’t have. You will only build trust if people believe what you’re telling them, and people can tell the difference between confidence and performance.

So start by declaring yourself both bankrupt and clueless. You genuinely don’t know this system yet, and you’re going to be humbled by it more than once, so don’t pretend otherwise. Say it plainly: this is a mess, I am not confident right now, and if I were a client of this service I would be frustrated too. You don’t need to be hysterical about it. You just need to be honest.

Trust gets built when people can reconcile what you’re telling them with what they’re already feeling. If the business has been living through outages and firefighting for months, and you turn up radiating false calm, you widen the gap between their experience and your narrative. That gap is where confidence goes to die. But if you can look frustrated, exhausted stakeholders in the eye and name what they’re feeling accurately, before you’ve fixed a single thing, you’ve already started earning the right to lead them out of it.

2. A Good Month Is Not Necessarily Good Risk Management

I would often arrive in an environment and see something strange happen at month end. If there had been no major incidents, the team would celebrate: “Great month. Everything was stable.” Then the next month there would be several outages, and suddenly the narrative would change: “Terrible month. Lots of incidents.” But very often almost nothing material had changed between those two months. The organization was effectively celebrating a winning scratch card. If you bought a scratch card and won R1,000, that does not mean you have developed an excellent investment strategy, you were lucky. Buying another one next month and losing does not mean your investment strategy suddenly deteriorated. Yet technology leaders do this all the time with operational risk: they narrate the outcome rather than manage the underlying probability.

The problem is that business stakeholders eventually notice. They hear technology declaring victory one month and explaining disaster the next, while being unable to explain what fundamentally changed between the two. Confidence disappears because the technology team is not actually managing risk, it is reporting the weather.

So I stopped celebrating clean months, and started celebrating changes in the underlying system: a vulnerability removed, an obsolete operating system retired, a database upgraded, a single point of failure eliminated, a recovery process automated, a capacity constraint understood, a dependency isolated, a rollback mechanism created. I would even celebrate a failure if it taught us something important that we could permanently eliminate, because a bad outcome that reduces future uncertainty can be more valuable than a good outcome that teaches you nothing.

3. When You Stop Understanding a System, Risk Turns Into Witchcraft

One of the most revealing characteristics of struggling technology organizations is superstition. Nobody explicitly calls it that, of course, it sounds more sophisticated. “We don’t deploy on Fridays.” “Don’t restart that server.” “That application doesn’t like being failed over.” “We only run that process at night.” “Don’t upgrade that library, because the last time we did, something went wrong.” Eventually the organization’s knowledge of cause and effect deteriorates so badly that operating the system starts to resemble witchcraft, please don’t wear red on Friday, because the last time somebody wore red on Friday we had an outage.

That happens because people no longer possess a sufficiently detailed causal model of the environment. They know that something once went wrong, but they do not really understand why, and the natural response is fear. Once fear takes over, the system begins to look like a giant Jenga tower, where every technology change feels equivalent to pulling out another block. So everyone agrees on what appears to be the safest strategy: don’t touch anything. Unfortunately, that is frequently the most dangerous strategy available.

It is also a vicious cycle. Lack of understanding produces unexplained failures. Unexplained failures produce folklore. Folklore produces fear. Fear produces avoidance. And avoidance produces even less understanding, because every change you avoid out of fear is also a change you never learn from. The knowledge that would have made the next intervention safer never gets acquired, precisely because nobody was willing to go and acquire it.

4. Gross Risk and Net Risk Are Completely Different Questions

This is where I think technology leaders frequently make a fundamental mistake: they look at gross risk. Imagine you have inherited an unstable twenty-year-old system containing obsolete infrastructure, unsupported software, fragile integrations, undocumented dependencies and manual recovery procedures. Its risk is already far beyond anything you would willingly design today. Now somebody proposes an upgrade, and the conversation immediately becomes: “Is the upgrade risky?” Of course it is. Almost everything you need to do is risky. But that is the wrong question. The right question is: if we make this change, are we increasing or decreasing the total risk of the environment? That is net risk thinking.

Risk disciplines already distinguish between the risk that exists before controls and the risk remaining once controls or responses have been applied, NIST, for example, defines residual risk as the risk remaining after controls or risk responses have been applied. Technology leaders need to apply the same logic dynamically. I might already have far more risk than I want, and I cannot magically move from enormous risk to zero risk. I have to transact my way down. Change A might introduce three units of execution risk while removing twenty units of structural risk. Change B might create ten units of risk and remove only one. They are not equivalent decisions simply because both involve “change”, the critical question is the direction of travel.

Does this move us toward a healthier system? Does it reduce an important structural risk? Will we learn something valuable? Can we limit the blast radius? Can we detect failure quickly? Can we roll it back? Is this a one-way door or a two-way door? Those questions allow you to manage a dangerous environment. “Everything is risky, therefore do nothing” does not.

To be precise about the term: I am using net risk here in the practical sense a leader actually needs, not as a formal taxonomy. It means the risk position you are left with after weighing both the execution risk an intervention introduces and the structural risk that intervention removes. That is the only comparison that matters when you are deciding whether to touch a fragile system.

5. The 200 Kilogram Man on the Sofa

Imagine a man who weighs 200 kilograms, spends every day sitting on a sofa and eats pizza continuously. His doctor tells him that his health is becoming extremely dangerous. He replies: “I can’t exercise. Exercise could cause a heart attack.” He isn’t completely wrong, suddenly attempting a marathon would indeed be dangerous. But concluding that he must therefore remain on the sofa eating pizza is absurd. The thing making intervention dangerous is also the thing that makes intervention necessary.

Legacy technology frequently works the same way. The system is so brittle that upgrading it is risky. The database is so old that migrating it is risky. The deployment process is so unreliable that releasing frequently is risky. The recovery mechanism is so poorly tested that deliberately exercising it feels risky. Therefore, the organization freezes everything, and another year passes. Now the operating system is older, the skills are rarer, the dependency chain is more complicated, the engineers who built it have left, and the next change is even more dangerous. You have not reduced the risk. You have compounded it.

6. A Change Freeze Can Produce a Very Misleading Number

Early in my career I saw a new CIO arrive and put an organization into roughly a 90 day change freeze. Incidents dropped, and this was celebrated as evidence that the strategy was working. But there is an obvious problem with that conclusion. It is like a football coach responding to a losing streak by refusing to play any more matches and then proudly announcing: “We haven’t lost a game in three months.” Correct. But you are no longer playing football.

A technology organization exists partly to change technology. It has to patch vulnerabilities, replace obsolete components, improve architecture, increase capacity, deliver products and respond to changing customer needs. Change is not an unfortunate side effect of technology management, change is part of technology management. The objective therefore cannot be to eliminate change. The objective is to become extremely good at changing things safely.

Research on high-performing software organizations supports this idea. DORA explicitly measures both software delivery throughput and stability, rather than treating them as opposing goals, and its current measures include deployment frequency and lead time alongside change failure rate, recovery time and deployment rework. It also recommends reducing batch size, because smaller changes are easier to understand and easier to recover from when something goes wrong. Perhaps even more interestingly, DORA’s research into change approval found no evidence that heavyweight external approval processes were associated with lower change failure rates, while such processes were associated with worse software delivery performance overall. Governance activity and risk reduction are not automatically the same thing.

Wrapping a problematic system in layers of governance solves nothing. It doesn’t make the system safer, and it doesn’t make the next change less dangerous. All it really does is announce to the organization that you are both clueless about the problem and without the skills to solve it, and that you’re hoping a committee can substitute for competence. It can’t. Approval boards don’t understand the database any better than you do.

7. The Answer to Dangerous Change Is Better Change

When I inherited environments like these, I tended to behave almost opposite to the prevailing culture. I wanted to exercise the system. I wanted proper nonproduction environments in which we could learn how it behaved, load tests, failure tests, recovery tests, database restoration tests, infrastructure changes, version upgrades. Sometimes I would deliberately perform changes at two in the morning, because I knew they contained risk and wanted maximum room to recover if that risk materialized. But the important part was not bravery. It was preparation.

These were not routine deployments, and I am not describing how ordinary software should ship. A healthy system should eventually make that kind of intervention unnecessary. This was how I handled unusually dangerous interventions into unhealthy systems, at a stage when the system itself had not yet earned the right to a calmer release process.

Before taking risk, I wanted to understand the rollback. How quickly can we detect that this has gone wrong? How do we restore the previous state? What telemetry tells us whether the change is healthy? How much of the estate should experience the change initially? What happens if our rollback itself fails?

This philosophy is visible in modern reliability engineering. Google’s SRE guidance describes canary releases as deliberately exposing only a small portion of production to a change, so that teams can learn from real traffic while limiting the amount of risk being consumed, smaller releases also make rollback easier. Chaos engineering takes the idea even further: its purpose is explicitly to experiment on systems in order to discover weaknesses and build confidence in their ability to survive turbulent conditions. The underlying principle is important. Confidence should come from evidence, not from avoiding the test.

8. Reliability Is Not the Absence of Change

Google’s Site Reliability Engineering work contains another useful concept: the error budget. The idea is remarkably simple. Rather than pretending that 100 percent reliability is achievable, teams establish an acceptable reliability target and then use actual system performance to determine how much operational risk they can currently tolerate. When reliability is comfortably inside the target, teams have room to make changes. When the budget is exhausted, effort shifts into testing and hardening the service rather than shipping new features.

That is very different from “we had an outage, therefore freeze everything.” It makes risk conditional and measurable, and most importantly, it acknowledges the tradeoff that technology leaders sometimes refuse to acknowledge: change creates risk, but not changing creates risk too. Security vulnerabilities accumulate. Technology becomes unsupported. People with critical knowledge leave. Capacity assumptions expire. Dependencies change underneath you. Disaster recovery processes decay because nobody exercises them. The gap between your current architecture and what you need becomes larger. Doing nothing is still a risk decision, it just happens to be one whose consequences are delayed.

But there is a sharper point hiding inside the error-budget idea, and it is worth pulling out explicitly, because the usual illustration only works cleanly when your own releases are the main source of unreliability. Stop shipping discretionary changes, and a system like that recovers. Legacy environments are frequently not that system. Imagine a platform that is burning its error budget because the database keeps hitting its limits, failover doesn’t work reliably, an obsolete operating system misbehaves on its own schedule, capacity is permanently marginal, recovery is manual, infrastructure fails without anyone touching a deployment pipeline, and the people who understood the dependencies have already left. You can freeze every application release tomorrow, and that system will keep consuming its error budget anyway. The budget is defined by actual service reliability, not merely by deployment-induced failure, so an exhausted budget doesn’t tell you to do less. Sometimes it tells you the opposite.

Sometimes leadership has to deliberately create room for a remediation effort, an unusually high rate of change over a defined period, aimed squarely at retiring the structural risk that is continuously eating the budget, even though some of that change will itself cause short-term disruption. That means going to customers and saying something leaders rarely say out loud: we have inherited, or allowed to develop, a system whose current reliability is unacceptable, fixing it requires more change than usual for a while, some of that change may cause further disruption, and the alternative, continuing on the present trajectory, carries a greater cumulative risk than the remediation programme does. That is what net risk thinking looks like when you apply it at the level of an entire remediation strategy rather than a single change. The remediation itself is gross risk. Whether it was the right call is determined entirely by how much structural exposure it retires.

There is a deeper version of this problem that most organizations already know about and rarely say out loud. A legacy platform can reach a point where its desired reliability target and its current engineering capability are simply incompatible. Customers know it, because they experience the failures. Engineers know it, because every change terrifies them. Executives know it, because the same platform keeps reappearing in incident reviews. And still everyone behaves as though holding the normal change constraints in place will eventually restore normal service. It won’t. At that point, leadership sometimes has to renegotiate reliability expectations temporarily in order to restore them permanently, which means saying the uncomfortable thing plainly: the safest course available to us still contains risk.

9. Everything Above Assumes Someone Still Understands the System

None of this works without a quiet precondition running underneath it: net risk thinking, error budgets, canaries, the whole apparatus, all depend on somebody in the organization still possessing a real causal model of how the system behaves. Take that away and every technique in this piece collapses back into gross risk thinking, because gross risk thinking is what’s left when nobody in the room can actually evaluate the trade.

This is the part that disappoints me most about the profession. Technology leaders give up on staying technical remarkably early, often well before they’ve earned the right to. And it rarely happens as a decision. Nobody wakes up and chooses to stop understanding the database. It happens through a thousand small substitutions: a stakeholder meeting instead of a postmortem, a status deck instead of a code review, a governance process instead of an afternoon spent reading telemetry. Each substitution looks reasonable in isolation. The traditional career ladder rewards every one of them, because the ladder only goes one way, up and out of the system, toward budgets, headcount and steering committees. Stay close to the machine too long and you start to look like you haven’t been promoted.

The cost of that ladder is that the person with the most authority to decide whether a change is safe is frequently the person furthest from being able to judge it. So they default to what they can still evaluate: process, optics, approvals, the appearance of control. Not because they’re incompetent, but because the organization spent years systematically pulling them away from the one skill that would have let them do better. They become mascots, pulled around by whatever the last incident or the last quiet month happened to suggest, because the currents of random chance are the only signal left available to someone who can no longer read the system directly.

The honest fix is structural, not motivational. Telling leaders to “stay technical” doesn’t survive contact with a career ladder that punishes them for it. What actually works is building a second ladder that doesn’t force the trade in the first place. This is the real argument for staff and principal engineer tracks, and for treating them as genuine peers to management roles rather than consolation prizes for people who “didn’t make it” into leadership. A principal engineer earns seniority, scope and influence precisely by going deeper into the system, not by leaving it. They keep reading the telemetry. They keep doing the 2am change. They keep the causal model alive, at exactly the level of seniority where an organization needs someone who can tell the difference between execution risk and structural risk, and where that judgment finally has the authority to matter.

Done properly, this doesn’t just preserve a pool of technical depth somewhere in the org chart. It changes who is in the room when net risk decisions actually get made. A principal engineer who has kept their hands on the system can sit across from a business stakeholder and say, with earned authority rather than borrowed confidence, “this specific change reduces our structural exposure more than it increases our execution risk, and here is how I know.” That sentence is only credible coming from someone who never stopped being able to evaluate it. A pure management track, however well-intentioned, cannot manufacture that credibility retroactively once the technical judgment has atrophied.

10. Celebrate Risk Retired, Not Luck Experienced

This changed how I thought about operational leadership. I stopped asking whether we had experienced a good month. I wanted to know whether we had built a better system. How many known risks disappeared? How much uncertainty did we remove? What did we learn? What can now fail without affecting customers? What can we now recover in minutes that previously took hours? What obsolete component no longer exists? What previously terrifying procedure has become routine? Those are much better indicators of progress than whether fortune happened to smile on you for thirty days.

A technology leader’s job is not to sit beside a fragile system hoping nothing happens. It is to progressively take control of the system: understand it, instrument it, exercise it, change it, learn from it, and systematically reduce the amount of risk embedded inside it. You will occasionally make mistakes while doing this, that is unavoidable. The objective is to make mistakes whose blast radius you can tolerate, whose consequences you can reverse, and whose lessons permanently improve the environment.

Because eventually every fragile technology organization faces the same choice. You can preserve the Jenga tower and spend your career begging everyone not to touch it. Or you can learn how the tower works and start rebuilding it. The first approach can give you some remarkably quiet months. The second is how you actually reduce risk.

Don’t be scared of the Jenga tower. Your job is to dominate it, to understand it, to subjugate it, to strip out the complexity that made it frightening in the first place. If you don’t have the skills, go find them. But never sit around hoping the technology will quietly get better on its own. It won’t.

References

NIST’s definition of residual risk provides the formal risk management analogue for the distinction between underlying exposure and what remains after controls are applied. NIST: Residual Risk

Google’s SRE work on error budgets explains why reliability and change velocity should be managed together rather than treating zero change as the definition of safety. Google SRE: Embracing Risk

DORA’s software delivery research measures throughput and stability together and advocates smaller, more recoverable changes. DORA Software Delivery Performance Metrics

Google’s guidance on canary releases describes how controlled production exposure can generate learning while limiting the cost of failure. Google SRE: Canarying Releases

The Principles of Chaos Engineering describe deliberate experimentation as a method for discovering weaknesses before they become uncontrolled production failures. Principles of Chaos Engineering