The Death of the Enterprise Warehouse (at Least for Product Reporting)
For product reporting, the enterprise warehouse is being replaced by domain read replicas. Each product team reports from a replica of its own operational database, getting accurate data seconds old instead of a day old. AI agents with read only access can then join data across several domain replicas when cross domain questions arise.
For the better part of two decades, almost every large organisation I have worked in or alongside followed the same ritual, and we followed it so faithfully that very few people stopped to ask whether it was still serving anyone. Every system of record, regardless of what it did or who owned it, was expected to shunt its product data into a central warehouse, and once the data arrived there it became somebody else’s problem, owned by a team that had never built the product, never spoken to its customers and never had to explain why a number looked strange on a Monday morning.
This post is about why I think that model is dying for product reporting, what replaces it, and where the warehouse genuinely still earns its keep.
1. The Ritual We All Performed
The pattern was remarkably consistent across industries. Product systems would extract their data on a schedule, usually overnight, and push it through a chain of ETL jobs into a central store, where a separate team would obsess over data quality rules, reconciliations and exception reports, trying to prove that what landed in the warehouse matched what existed in the source. Because the data had been flattened, renamed and reshaped along the way, a large part of that effort was spent rediscovering meaning that the product team had known all along and that the pipeline had quietly thrown away.
Once the data was in, we hired teams of data scientists and analysts to report on product specific questions from the warehouse, which meant that the people closest to the data were no longer the people answering questions about it. When the queries those analysts wrote became slow, which they inevitably did because a warehouse serving every domain cannot be tuned for any single one of them, we hired warehouse engineers whose job was to optimise the queries going into the warehouse, build aggregate tables, manage partitions and negotiate compute budgets between competing teams.
Each step made sense locally, and each step added another layer of people between a question and its answer.
2. The Dashboard Graveyard
The visible output of all this investment was dashboards, and we produced them in staggering quantities. Most large enterprises I know have thousands of them, and if you look at the access logs honestly, a large proportion are opened rarely or never, many contradict each other because they were built on slightly different definitions of the same metric, and almost all of them are at least a day out of date because they sit at the end of an overnight batch.
A dashboard that is twenty four hours stale is a strange thing to run a digital product on. If a release goes out at ten in the morning and breaks a conversion funnel, the warehouse will tell you about it tomorrow, by which point your customers, your call centre and probably social media have already told you. The product team ends up building its own operational views anyway, because it has to, and the warehouse dashboard becomes a historical artefact that someone presents in a monthly meeting.
None of this is a criticism of the people involved, many of whom are excellent. It is a criticism of an architecture that separated the ownership of data from the understanding of it, and then tried to fix the consequences with headcount.
3. What Changed: Every Domain Already Has the Data
The question I keep coming back to is simple: why can’t product teams report on their own domain directly from a read replica of their own database?
The data in a product’s operational store is the most accurate representation of that product that exists anywhere in the organisation, because it is the data the product actually runs on. It has not been filtered, transformed or reinterpreted by a pipeline built by someone else, and it reflects the product’s real semantics, including all the awkward edge cases that tend to get smoothed away on the journey to a central store.
A read replica gives you that data without putting load on the primary, and with replication lag that is typically measured in seconds or less rather than hours. It is not the same as reading the primary, and I will come back to the consistency caveats, but the difference between “a few seconds behind” and “yesterday” is the difference between operating a product and reading its obituary.
There are several practical benefits that fall out of this almost for free. The access patterns on a replica can be tuned for the reporting workload of that one domain, with indexes and materialised views designed around the questions that domain actually asks, rather than compromising across every business unit in the company. Permissions also become dramatically simpler, because you are applying access control at the boundary of a single domain whose owners understand exactly what is sensitive, rather than trying to retrofit row and column level security onto a giant shared schema where nobody is quite sure who should see what. In my experience the permission model of a central warehouse is one of the quietest and largest sources of audit pain in an enterprise, and shrinking the blast radius of each grant is a real win.
At Capitec we have deliberately engineered the stack for real time agentic workloads using read replicas, and the operational benefits for reporting have been one of the more pleasant side effects of that decision.
4. AI Joins Across Domains
The traditional objection to domain reporting is the cross domain question. If I want to know how a change in the credit product affected card spend for customers who also use a particular savings feature, I need data from three domains, and the warehouse was the only place where all three lived side by side.
This is where AI changes the economics. An agent with read only access to several domain replicas, together with a good description of each domain’s schema and semantics, can plan and execute those cross domain queries on demand, pulling the relevant slices from each replica and joining them in the context of the specific question being asked. You no longer need to pre build and maintain a single conformed model of the entire enterprise just in case someone asks a question that spans it; you need well described domains and a reasoning layer that can stitch them together when someone does.
I want to be careful about the claim here, because it is easy to overstate. The AI is not magic, and the quality of its joins depends heavily on how well each domain documents its entities, keys and definitions. If two domains disagree about what a “customer” or an “active account” is, the agent will either surface that disagreement or, worse, paper over it, so those shared definitions still need to exist and be owned. What changes is that the definitions live as contracts at the domain boundary rather than as transformations buried inside a pipeline, and the expensive, brittle work of physically moving and reshaping everything into one place largely goes away.
5. The Honest Caveats
I am arguing for a shift, not pretending that read replicas solve every problem, so it is worth being explicit about where this approach needs care.
Replica consistency is the first one. An asynchronous replica is eventually consistent with its primary, so a report can be a few seconds behind, and under heavy write load or during an incident that lag can grow. Within a single domain, reading from one replica gives you a coherent view of that domain as of a point in time, which is more than most overnight batches can honestly claim. Across domains, however, an agent querying several replicas is reading each at a slightly different moment, so cross domain answers are near real time rather than a single transactionally consistent snapshot. For product decisions that is almost always fine; for anything that needs to reconcile to the cent at a fixed cutoff, it is not, which leads directly to the next section.
The second caveat is schema discipline. When reporting reads directly from a domain’s store, that schema becomes a contract, and product teams need to treat breaking changes to it with the same seriousness they would apply to a breaking change in a public API. This is a cultural shift as much as a technical one, but it is a healthier one than the current situation, where schema changes silently break pipelines that the product team did not know existed.
The third is governance of sensitive data. Giving an AI agent access to multiple domains means you need to think carefully about what it is allowed to combine, how its queries are logged, and how personal information is masked before it leaves a domain. The good news is that the narrower, domain scoped permission model described above makes this considerably more tractable than it is in a single warehouse where everything is already sitting together.
6. Where the Warehouse Still Belongs
None of this means you should switch off your warehouse tomorrow. There is a genuine and important place for centrally curated, historically stable data, and in banking the clearest example is regulatory reporting.
Regulatory submissions need fixed cutoffs, full lineage, reproducibility months or years later, and numbers that reconcile exactly across domains at a specific point in time. They need the ability to restate a period and explain precisely what changed. Those are exactly the properties a well run warehouse provides and a set of live replicas does not, and the same logic applies to statutory financial reporting, long horizon historical analysis and model training datasets that need to be frozen and versioned.
The mistake was never building warehouses. The mistake was making the warehouse the default destination for every question, including the overwhelming majority of product questions that care far more about accuracy and freshness than about a reconciled month end snapshot.
7. One Engineer per Domain
The part of this that I find most compelling is what it does to the cost of answering questions. In the old model, getting a new product metric in front of a decision maker could involve a source system team, an ETL team, a data quality team, a warehouse engineering team and an analytics team, each with their own backlog and their own priorities. It is not unusual for a reasonable question to take weeks.
In the domain model, a single engineer who understands the product can stand up a replica, tune it for the domain’s reporting needs, expose well described views, and answer most of the product’s questions directly, while the AI layer handles the cross domain questions that would previously have required a central team. That engineer is not doing heroic work; they are simply not paying the coordination tax that the warehouse architecture imposed on every question.
That is the real point for me. The enterprise warehouse for product reporting is not dying because the technology stopped working; it is dying because the organisational overhead it created is no longer necessary, and once teams experience accurate, unfiltered, near real time data in their own domain, very few of them want to go back to waiting until tomorrow to find out what happened today.