The Death of the Enterprise Warehouse (at Least for Product Reporting)

The Death of the Enterprise Warehouse (as the Mandatory Route for Reporting)

👁5views

For product reporting, the enterprise warehouse is being replaced by domain read replicas. Each product team reports from a replica of its own operational database, getting accurate data seconds old instead of a day old. AI agents with read only access can then join data across several domain replicas when cross domain questions arise.

Article Summary
  • 1.
    What it is
    Product reporting from a central enterprise data warehouse is a model the author argues is dying. The article explains why product teams can report directly from read replicas of their own databases, and how AI agents handle cross domain questions.
  • 2.
    Why it matters
    Reporting from a domain's own replica gives data that is seconds old instead of a day old, keeps ownership with the people who understand the data, and simplifies permissions and tuning for a single domain.
  • 3.
    Key takeaway
    The difference between a replica a few seconds behind and a warehouse that is a day behind is the difference between operating a product and reading its obituary.
~12 min read

For the better part of two decades, almost every large organisation I have worked in or alongside followed the same ritual, and we followed it so faithfully that very few people stopped to ask whether it was still serving anyone. Every system of record, regardless of what it did or who owned it, was expected to shunt its product data into a central warehouse, and once the data arrived there it effectively became somebody else’s problem, owned by a team that had never built the product, never spoken to its customers and never had to explain why a number looked strange on a Monday morning.

This post is not an argument that warehouses are bad technology. Modern analytical platforms are remarkable pieces of engineering. It is an argument that the central warehouse should lose its status as the mandatory route for product reporting, and that each domain should own its reporting and choose the infrastructure its workload actually needs, whether that is a read replica, a domain owned analytical store or, for some questions, the warehouse itself.

1. The Ritual We All Performed

In the implementations I have encountered, the pattern was remarkably consistent. Product systems extracted their data on a schedule, usually overnight, and pushed it through a chain of ETL jobs into a central store, where a separate team obsessed over data quality rules, reconciliations and exception reports, trying to prove that what landed in the warehouse matched what existed in the source. Because the data had been flattened, renamed and reshaped along the way, a large part of that effort was spent rediscovering meaning that the product team had known all along and that the pipeline had quietly thrown away.

Once the data was in, we hired teams of data scientists and analysts to report on product specific questions from the warehouse, which meant that the people closest to the data were no longer the people answering questions about it. When their queries slowed down, we hired warehouse engineers to optimise the queries going into the warehouse, build aggregate tables, manage partitions and negotiate compute budgets between competing teams.

It is worth being fair here: none of that was forced on us by the technology. Platforms like Snowflake let you run separate compute for separate workloads, and features such as dynamic tables can target freshness of around a minute, even if actual lag can exceed the target and nothing about them guarantees that the source data arrives promptly. Daily batches and shared compute contention were choices, not laws of physics. But the organisational model survived even where the technology improved, and that is the part I want to question.

2. The Dashboard Graveyard

The visible output of all this investment was dashboards, and we produced them in staggering quantities. In the enterprises I know well there are thousands of them, and if you look at the access logs honestly, a large proportion are opened rarely or never, many contradict each other because they were built on slightly different definitions of the same metric, and most of them sit at the end of an overnight batch.

A dashboard that is twenty four hours stale is a strange thing to run a digital product on. If a release goes out at ten in the morning and breaks a conversion funnel, a nightly pipeline will tell you about it tomorrow, by which point your customers, your call centre and probably social media have already told you. The product team ends up building its own operational views anyway, because it has to, and the warehouse dashboard becomes a historical artefact that someone presents in a monthly meeting, reading the product’s obituary rather than helping to operate it.

None of this is a criticism of the people involved, many of whom are excellent. Even against a fast, well configured warehouse, the core problem remains: a product team should not have to transfer ownership of its data or join another team’s queue in order to answer a question about its own product.

3. Domains Owning Their Reporting

The question I keep coming back to is simple: why can’t product teams report on their own domain, using the data their product actually runs on, without routing every question through a central team?

The data in a product’s operational store is the most accurate representation of that product’s current state that exists anywhere in the organisation. It has not been filtered or reinterpreted by a pipeline someone else built, and it reflects the product’s real semantics, including the awkward edge cases that tend to get smoothed away on the journey to a central store. The question is how to report on it without hurting the product, and there are really two building blocks, which are easy to conflate but behave quite differently.

The first is a physical read replica, such as a PostgreSQL hot standby or an Aurora reader. This gives you separate compute for reporting queries, with replication lag that is usually measured in seconds or less, but it uses exactly the same schema and indexes as the primary. You cannot create reporting specific indexes on a PostgreSQL standby, and an Aurora reader shares the cluster’s storage volume with the writer. Any reporting index or materialised view you want has to be created on the primary, where it carries its maintenance cost on every write. A replica therefore offloads reporting compute rather than providing complete independence: on PostgreSQL, long standby queries can conflict with replication replay and get cancelled, and if you enable hot_standby_feedback to avoid that, you can delay vacuum cleanup on the primary instead. These are manageable engineering trade offs, but they are trade offs.

The second is a separate reporting store owned by the domain and fed through logical replication or change data capture. This is where you get genuine freedom to design indexes, transformations and reporting structures around the questions your domain actually asks, without imposing any of that cost on the operational database. It is a little more infrastructure, but it remains inside the domain’s ownership, built by people who understand what the data means.

Most domains will want both: a replica for lightweight, near real time operational questions, and a domain owned analytical store for anything heavier. At Capitec we have deliberately engineered the stack for real time agentic workloads using read replicas, and the reporting benefits have been one of the more pleasant side effects of that decision. The point is not that the replica answers everything; it is that the domain decides.

4. Product Reporting Needs History Too

There is a qualification here that is easy to miss and that matters a great deal. A great many ordinary product questions depend on history rather than current state: conversion before and after a release, retention by onboarding cohort, or what status an account had at the moment a customer was shown an offer. An operational database typically stores the current version of each record, so the earlier states those questions depend on may simply not exist in it.

My own funnel example illustrates the problem. A customer who abandons a journey halfway through may never create a record in the product database at all, so no replica, however fresh, will ever reveal that the abandonment happened. The only way to answer that question is to capture the events in the first place.

So the model I am describing explicitly includes domain owned event history alongside the operational data: the domain emits and retains the events and state changes it cares about, in a store it controls, and reports on them directly. That is not a return to the central warehouse; it is the domain taking responsibility for the history of its own product rather than outsourcing it to a pipeline it does not understand.

5. Publish an Interface, Not Your Tables

Once reporting reads from a domain’s data, it is tempting to say that the operational schema becomes a contract that must be treated like a public API. That would be a mistake, because it recreates the very coupling we are trying to remove: every internal table change would have to be negotiated with every reporting consumer.

The better pattern is for each domain to publish stable, versioned reporting views or interfaces, with clear metric definitions and documented semantics, and to treat those as the contract. Internal tables can then evolve as the product needs them to, while the domain remains accountable for the metrics it publishes. This is also what makes the next step possible, because a well described interface is exactly what an AI agent needs in order to reason across domains.

6. AI Across Domains

The traditional objection to domain reporting is the cross domain question. If I want to know how a change in the credit product affected card spend for customers who also use a particular savings feature, I need data from three domains, and the warehouse used to be the only place where all three lived side by side.

AI changes the effort involved in asking that question, but it is worth being precise about what the AI actually does, because the model is not joining large datasets inside its context window. In the architecture I have in mind, the agent interprets the question, works out which domain interfaces are relevant and proposes a query plan. Governed database or analytical tools then execute the joins and calculations, with filtering and aggregation pushed as close to each source as possible and sensible limits on what any single query may scan or move. The answer comes back carrying the queries that produced it, the metric definitions used, the timestamps of each source and any checks that were run, so that a human can see exactly how the number was derived.

When a question would require scanning and moving millions of rows, or when the same cross domain question keeps being asked, the right response is to materialise a governed analytical dataset for it rather than recomputing it on demand. In other words, AI removes much of the human coordination involved in asking cross domain questions, but the execution cost and the semantic obligations still need an architecture. If two domains disagree about what a “customer” or an “active account” is, the agent will either surface that disagreement or, worse, paper over it, so shared definitions still need to exist and be owned.

7. The Honest Caveats on Consistency and Governance

Consistency is the first caveat. An asynchronous replica is eventually consistent with its primary, so a report can be a few seconds behind, and under heavy write load or during an incident that lag can grow. Within a single domain, you can get a coherent point in time view, but only if the report runs inside an appropriate transaction: under PostgreSQL’s default Read Committed isolation, each statement sees its own snapshot, so a report assembled from several queries needs something like Repeatable Read when one consistent snapshot matters. Across domains, an agent querying several stores reads each at a slightly different moment, so cross domain answers are near real time rather than a single transactionally consistent snapshot. For most product decisions that is entirely acceptable; for anything that must reconcile to the cent at a fixed cutoff, it is not.

Governance is the second. Domain ownership can make access decisions clearer, because the people granting access understand what is sensitive, but simpler permissions are an outcome you have to engineer rather than something you get for free. An agent that is allowed to read several domains can produce combinations that none of the individual grants anticipated, so you need controls over the combined result, query logging, and masking of personal information before it leaves a domain, not merely permission to read each source.

8. Where the Warehouse Still Belongs

None of this means you should switch off your warehouse tomorrow. There is a genuine and important place for centrally curated, historically stable data, and in banking the clearest example is regulatory reporting.

Regulatory submissions need fixed cutoffs, full lineage, reproducibility months or years later, and numbers that reconcile exactly across domains at a specific point in time. They need the ability to restate a period and explain precisely what changed. Those are exactly the properties a well run warehouse provides and a collection of live domain stores does not, and the same logic applies to statutory financial reporting, long horizon enterprise analysis and model training datasets that need to be frozen and versioned.

The mistake was never building warehouses. The mistake was making the warehouse the default and often mandatory destination for every question, including the large majority of product questions that care far more about accuracy, freshness and ownership than about a reconciled month end snapshot.

9. What One Engineer Can Do

The part of this I find most compelling is what it does to the cost of answering questions. In the old model, getting a new product metric in front of a decision maker could involve a source system team, an ETL team, a data quality team, a warehouse engineering team and an analytics team, each with its own backlog and priorities, and it was not unusual for a reasonable question to take weeks.

This is not a universal staffing formula, but with a shared platform, established controls and a bounded reporting workload, a single engineer who understands the product can go a remarkably long way: provisioning a replica and a domain analytical store, capturing the events the product needs, publishing well described reporting interfaces and answering most of the domain’s questions directly, while the AI layer handles cross domain questions that would previously have required a central team. That engineer is not doing heroic work; they are simply not paying the coordination tax that the central warehouse model imposed on every question.

That is the real point for me. The enterprise warehouse is not dying because the technology stopped working. It is losing its claim to be the mandatory route for product reporting because the organisational overhead it created is no longer necessary, and once teams experience accurate, near real time data in their own domain, very few of them want to go back to waiting until tomorrow to find out what happened today.

Leave a comment

Your email address will not be published. Your first comment is held for approval.