The Futility of Corporate Metrics: Why Measurement Isn't Movement

Measurement Isn’t Movement: When Corporate Metrics Become a Substitute for Management and Execution

👁30views
Article Summary
  • 1.
    What it is
    Measurement is not movement: the article argues that ROI calculations, cost attribution and dashboards can become a substitute for useful action. It uses a bank mortgage book, branch networks and AI spending as examples.
  • 2.
    Why it matters
    It argues that once a strategic decision is made, attention should shift to making the capability work well rather than repeatedly justifying it with isolated ROI calculations.
  • 3.
    Key takeaway
    Measurement earns its place only when it leads to action, as measuring token usage and waste led to Claude Burst, expected to roughly halve the AI bill.
~17 min read
Listen to this article1 play

I recently saw an advert suggesting that approving an AI budget without ROI metrics was the equivalent of driving blind. It is a neat line, and it captures a belief that runs through a great many companies: if we can attribute enough costs, calculate enough returns and produce enough dashboards, we will make better decisions. What the advert leaves out is how easily all of that activity can quietly become a substitute for doing something useful, and how often the people producing the numbers are the same people who could have been fixing the problem the numbers describe.

Here is a small provocation to keep in mind as you read. A few weeks ago one of our clients left a four star review of our app, called it the best banking app in South Africa by far, and then told us in two sentences exactly what we would need to build to earn the fifth star, and why it mattered to them. No ROI model would have found that, no cost allocation would have surfaced it and no dashboard would have turned red, because nothing was broken. The most valuable piece of product information we received that month arrived unprompted, unformatted and free, and the only way it creates any value is if someone reads it and goes and builds the thing. Hold that thought, because I will come back to it.

I am not arguing against measurement. A business that does not understand its costs, its risks and the quality of its service is in trouble. My argument is against the growing habit of treating more granular measurement as more sophisticated management, against the time, attention and organisational energy that habit consumes, and against our natural preference for the data that reassures us over the data that asks us to act.

1. The Numbers You Already Have

A company already has financial measures that matter. It knows its revenue, its costs, its profitability and the capital it consumes, and those numbers are audited, reconciled and broadly trusted. Breaking those measures down further can answer genuinely important questions about products, customers and operations, but increasing the level of detail does not automatically increase the level of understanding, and past a certain point it can reduce it, because the allocations start to carry more assumptions than facts.

W. Edwards Deming listed “management by use only of visible figures” among his seven deadly diseases of management, pointing out that many of the figures that matter most to a business are unknown or unknowable, such as the long term cost of an unhappy customer or the value of a happy one [1]. That is not an argument for ignoring the visible figures. It is a reminder that the spreadsheet is a partial picture of the business, and that the parts it leaves out are often the parts that decide whether the business thrives.

2. The Mortgage That Looks Like a Bad Idea

Consider a bank’s mortgage business. Viewed in isolation, a mortgage book can offer modest returns while carrying substantial funding requirements, meaningful operational overheads and real credit risk. A product level income statement might reasonably suggest that the bank would be better off putting its capital and its effort somewhere else, and on its own terms that analysis may be perfectly correct.

The difficulty is that clients do not experience a bank as a collection of independent product income statements. They experience one relationship, and buying a home is one of the largest financial decisions most people ever make. If you cannot help them do it, there is a real risk they move to a bank that can, and when they go they may take their transactional account, their savings and their other profitable business with them. The academic evidence on how strongly banking relationships bind products together is instructive here. Basten and Juelsrud, studying household banking data and publishing in the Review of Financial Studies, found that a bank was 20 percentage points more likely to sell a loan to an existing depositor than to an otherwise comparable household, and that this effect appeared to be driven mainly by customer stickiness rather than by any information advantage the bank had about those customers [2]. Their study looks at deposits leading to loans rather than at what happens when a bank withdraws a product, so it should not be read as proving that dropping mortgages would drive clients away. What it does suggest is that value in retail banking tends to sit in the relationship, and that a product’s contribution can include something its individual return struggles to capture: keeping the wider relationship intact.

That does not mean every weak product deserves a strategic excuse, because plenty of underperforming products are simply underperforming. It means you need to understand how the parts of the business relate to one another before deciding what each part is worth. You can calculate a product’s profitability to several decimal places and still misunderstand its contribution to the company, since precision in the calculation does nothing to guarantee wisdom in the decision.

3. Once You Have Decided, Get Good At It

The same principle applies to capabilities you have already decided your business needs. If your clients need a branch network, then run an efficient branch network: understand demand, put branches in the right places, equip the people who work there properly and go actively looking for waste. Those are management problems, and they reward attention.

You should still challenge how you deliver that capability and revisit the underlying decision when circumstances genuinely change. It is also worth being precise about what the strategic decision actually covers. Deciding that clients need physical service does not establish that every existing branch, every location or the current operating model remains justified, and “strategic necessity” should never become protection for a particular implementation. The capability is settled; its footprint is permanently open to improvement. What does become a distraction is repeatedly asking the capability itself to justify its existence through an isolated ROI calculation, when the energy would be better spent making it work well.

A branch may support clients who generate most of their revenue elsewhere in the bank, through the app, through lending or through products sold years later. Trying to assign an exact share of that revenue to the branch quickly becomes an elaborate accounting exercise that does very little to improve either the client experience or the cost of serving those clients. You do need to understand the branch’s costs, service levels and usage, because those measures lead directly to practical improvements in where branches sit, how they are staffed and what they are for.

4. The Allocation Trap

The problem gets considerably worse when companies try to attribute every shared cost and every benefit to individual teams. Who pays for the platform, who gets credit for the improvement, and how much of an outcome belongs to engineering rather than to product, operations or distribution? These questions can consume weeks and sometimes months, particularly when the answers affect budgets and how people are judged, because at that point everyone involved has a perfectly rational reason to argue. The economist Charles Goodhart observed that a statistical regularity tends to collapse once it is used for control, an idea Marilyn Strathern later generalised into the familiar form that when a measure becomes a target, it ceases to be a good measure [3]. Allocation numbers that decide budgets are targets, and they behave accordingly.

The history of activity based costing is a useful cautionary tale. Robert Kaplan, who helped create the method, later wrote with Steven Anderson that many organisations had struggled with traditional activity based costing because of the expense of interviewing and surveying staff, the use of subjective time allocations that were costly to validate, and the difficulty of keeping the model current as processes, products and customers changed [4]. In one of the cases they describe, a company employed 14 people full time simply to maintain its costing model, and they note that managers could end up spending their time disputing the accuracy of the allocations rather than acting on them [5]. Their response was to simplify the approach dramatically, which is telling in itself: the people who built the more granular model concluded that its granularity was getting in the way.

Eventually, in most organisations, an allocation is agreed and a dashboard is updated. Meanwhile no client’s problem has been solved, no process has become faster and no unnecessary cost has been removed. The company has become more elaborate in its description of the work without becoming any better at doing it. Measurement can help us choose where to move and tell us whether we are making progress, but it cannot create that progress for us.

5. How We Approached AI

Our approach to AI at Capitec is a practical example of the alternative. We could have spent our time asking every team to prove its individual return before giving it access, building business cases and debating whose work justified the expenditure. We could have decided that some teams did not “deserve” AI because their contribution was difficult to translate into a tidy financial model.

Instead, we chose to use it at scale, learn from the work and optimise what we found. On the cost side, we focused our measurement on token usage and waste, because those were things we could understand and act on directly. That investigation led us to build Claude Burst, a local gateway that routes inference between a bounded subscription and metered API capacity, which we expect will roughly halve our AI bill. That is an expected saving rather than a realised one, and the measurement earned its place because it led directly to an engineering change.

A sceptical reader could fairly point out that this only proves we made AI cheaper, not that AI made the work better, and that cheapening an unproductive activity is no achievement. That is the right challenge, and it is why cost was never the only thing we looked at. The evidence that matters for productivity is practical and close to the work: whether things get finished faster, whether there is less rework, whether incidents are diagnosed and resolved sooner, and whether the output is better than it was. Those signals can be observed directly by the people doing the work and the leaders working alongside them, without requiring every team to translate each improvement into an attributed financial return. The discipline is in actually looking, and in being willing to stop using the tools where they are not helping.

If we had started by rationing access, we might well have produced a smaller bill, simply because fewer people were using the tools. We would also have had far less opportunity to learn where the tools helped, where they failed and where the waste actually sat, and a lower bill on its own would not have told us whether we had made the company any more productive.

6. Why Early Business Cases Mislead

There is a particular weakness in demanding a precise business case for a capability before people have had a reasonable opportunity to learn how to use it. At that stage you may be measuring their ability to write a persuasive forecast more than the tool’s ability to improve their work. The teams with easily counted outputs get access, while teams doing valuable but less easily attributed work struggle to justify themselves. Steven Kerr described this pattern in 1975 in a paper with the wonderful title “On the folly of rewarding A, while hoping for B” [6], and it plays out in budget committees to this day.

There is also good economic reason to be careful with early measurements of a general purpose technology. Brynjolfsson, Rock and Syverson describe what they call the productivity J curve: technologies like AI require large complementary investments in new processes, skills and ways of working, those investments are mostly intangible and poorly captured in the accounts, and so measured productivity growth tends to be understated in the early years and then overstated later, once the benefits of those earlier investments are being harvested [7]. Their work concerns national productivity statistics rather than individual companies, and it is an argument about measurement error, not a promise that any particular AI investment will pay off. The relevant point for a firm is narrower: the learning is itself an investment that standard measures struggle to see, so a demand that it pay for itself before it has happened will tend to produce the conclusion that it should not happen.

Spending still needs boundaries, and results still need scrutiny. Broad access does not remove the obligation to manage costs, examine quality or stop activities that are clearly going nowhere. Those decisions simply become better informed when they are grounded in actual use and direct observation rather than in forecasts written before anyone had tried the thing.

7. Data That Finds You and Data You Have to Find

We live in the information age, and the problem facing most leaders is not a shortage of data but an abundance of it. It helps to notice that this data arrives in two very different forms. The first kind lands on your desk without any effort on your part: the monthly pack, the scheduled dashboard, the green status report. Much of it is reassuring, some of it is flattering, and almost none of it asks anything of you. You read it, you feel that the business is under control, and you move on to the next meeting.

The second kind has to be hunted. It sits in client complaints, in app store reviews, in the workaround a consultant has built because the system does not quite do what clients need, and in the queue that forms every month end. It rarely arrives formatted, it is often uncomfortable to read, and when you find it, it is only valuable if you actually do something with it. That last part is the real difference between the two. Passive data asks for your attention; hunted data asks for your effort, and most people, given the choice, will prefer the data that calms them to the data that creates work for them. It is worth asking yourself honestly which of the two is more likely to build a better company.

Go back to the app review I mentioned at the start. What the client asked for was to select a transaction and see the full history of payments to that merchant over a year or longer, and the ability to generate a statement showing only money in or only money out over twelve or thirty six months, because it would make their tax return far easier. That is a precise, actionable product specification, written for free by someone who uses the product, along with the reason it matters to them. No ROI model would have surfaced it, no attribution exercise would have found it and no dashboard would have flagged it, because nothing is broken and the client is happy. It only becomes valuable if someone reads it, recognises how many other clients probably share the same need at tax time, and goes and builds it. Thanking the client for the suggestion is courteous; shipping the feature is what actually improves the company.

That is the pattern this whole argument rests on. The most useful information in a business is usually not the information that arrives on a schedule, and the measure of a leader is not how much data they consume but how much of it they turn into change.

8. Get Closer to the Work

Do not expect Power BI reports to run your business or make you wise. A dashboard is very good at showing you where to ask questions, but it cannot replace spending time in a branch, watching the queues form, understanding the exceptions and asking employees what stops them from helping clients. The Toyota tradition has a name for this, genchi genbutsu, roughly meaning going to the actual place to see the actual thing, and Taiichi Ohno was known for making managers stand on the factory floor and watch until they understood what was happening [8].

For productivity, my instinct is to get on with the work, do enough of it to learn, and pay close attention to the opportunities for improvement that appear along the way. Spend time seeing what people actually do, where they wait, what they repeat and which decisions they cannot make without permission. Watch someone move the same information between three systems, or spend an afternoon assembling a report that nobody acts on, and the opportunity becomes far easier to understand than any slide could make it. Give people useful tools, help them apply those tools wisely, and remove the obstacles you discover alongside them.

A dashboard might show that a team’s output is low. Working alongside that team may reveal that its members spend half their time resolving failures elsewhere in the organisation, protecting everyone else’s productivity at the expense of their own reported numbers. Judging them through a narrow measure can lead you to penalise precisely the behaviour the company most needs. The historian Jerry Muller calls the underlying pattern “metric fixation”, the belief that measuring, publishing and rewarding performance will by themselves produce improvement, and he documents its side effects across hospitals, schools, police forces and businesses while arguing that metrics work best as a complement to judgement rather than a replacement for it [9].

9. A Better Test for Any Measure

A useful test for any measure is to ask what it is for. Some measures inform a decision: they help us remove waste, improve quality, redirect investment or solve a client problem. Some monitor a risk: they tell us that a service remains within acceptable limits, or warn us early when it starts to deteriorate, and they are valuable precisely on the days when nothing needs to change. Some satisfy an obligation to a regulator, an auditor or a client. A measure that does one of those things, has an owner, and has an agreed response when it moves deserves the effort it takes to produce. The ones worth questioning are the reports with no identifiable purpose, no owner and no response when the number changes, and the allocation exercises whose main output is an argument.

Even a well chosen measure only earns its keep when it leads to something better. Taking action is not the goal in itself, because an intervention still has to produce an improvement, and the measure is how you find out whether it did. So measure the performance of the whole business, investigate the detail wherever it gives you something useful to act on or to watch, and be as efficient with the time and attention spent proving where value belongs as you are with money. Then go and work with your teams: equip them, direct them and help them solve the problems they actually have.

The company improves through the changes you make. Measure enough to guide those changes and check their effect, then keep moving. A more detailed account of standing still is still standing still.

References

  1. Deming, W. E. (1986). Out of the Crisis. MIT Center for Advanced Engineering Study.
  2. Basten, C. and Juelsrud, R. “Cross-selling in bank-household relationships: Mechanisms and implications for pricing.” The Review of Financial Studies, 39(6), 1751 to 1784. https://www.sfi.ch/en/publications/cross-selling-in-bank-household-relationships-mechanisms-and-implications-for-pricing
  3. Goodhart, C. A. E. (1975). “Problems of Monetary Management: The U.K. Experience.” Papers in Monetary Economics, Reserve Bank of Australia; and Strathern, M. (1997). “‘Improving ratings’: audit in the British University system.” European Review, 5(3), 305 to 321.
  4. Kaplan, R. S. and Anderson, S. R. (2003). “Time-Driven Activity-Based Costing.” SSRN working paper. https://papers.ssrn.com/abstract=485443
  5. Kaplan, R. S. and Anderson, S. R. (2004). “Time-Driven Activity-Based Costing.” Harvard Business Review, November 2004, 131 to 138.
  6. Kerr, S. (1975). “On the folly of rewarding A, while hoping for B.” Academy of Management Journal, 18(4), 769 to 783.
  7. Brynjolfsson, E., Rock, D. and Syverson, C. (2021). “The Productivity J-Curve: How Intangibles Complement General Purpose Technologies.” American Economic Journal: Macroeconomics, 13(1), 333 to 372. https://www.aeaweb.org/doi/10.1257/mac.20180386
  8. Ohno, T. (1988). Toyota Production System: Beyond Large-Scale Production. Productivity Press.
  9. Muller, J. Z. (2018). The Tyranny of Metrics. Princeton University Press. https://press.princeton.edu/books/hardcover/9780691174952/the-tyranny-of-metrics

Leave a comment

Your email address will not be published. Your first comment is held for approval.