Prototype Bloat: When AI Makes It Easier to Build Than to Decide
AI makes building software far cheaper, but it does not make deciding what to keep, own and support any cheaper. Organisations end up with many prototypes and few products, which the article calls prototype bloat. The old cost of building used to filter weak ideas, and AI removes that filter.
AI has collapsed the cost of turning an idea into working software. It has not collapsed the cost of deciding which of those ideas deserve to survive, and that gap is where a new kind of organisational waste is starting to accumulate.
1. Easier to build than to decide
AI has made it easier to build software than to decide whether that software should exist, and I think that is one of the most remarkable shifts in the economics of technology in my career. For decades the cost of engineering acted as a natural filter on ideas: we could always imagine far more than we could afford to build, so organisations were forced to prioritise. AI is removing that constraint. A product manager, designer or business analyst can now turn an idea into a convincing working application in hours, often without involving an engineering team at all.
There is a diagram doing the rounds that captures this well. On the “before AI” side, a small group of people generate ideas, a larger group turns them into working solutions and a much larger audience uses what gets built. On the “after AI” side the shape inverts: almost everyone can have an idea, almost everyone can build something that demonstrates it, and the number of people using any given thing appears to shrink. It is a provocative illustration rather than a measured prediction, but it points at a risk that most technology leaders are already starting to feel, which is that the supply of software may grow enormously while the demand for any individual piece of it does not.
This is enormously valuable, and I want to be clear from the outset that the answer is not less prototyping, because a working prototype may be the best specification tool we have ever had. But it creates a new organisational problem: when everyone can build, who decides what should survive? I call the result prototype bloat, which I would define as the accumulation of software experiments that outlive the decisions they were created to inform. The problem is not that we are creating too many prototypes. It is that our ability to create them is growing much faster than our ability to compare, consolidate and retire them. AI has made divergence cheap, and the leadership challenge now is to make convergence effective.
A word on evidence before going further. Prototype bloat is a new enough phenomenon that nobody has measured it directly, so none of the research I cite below proves it exists. I use three kinds of evidence instead, and I try to be clear which is which: studies of mechanisms that plausibly drive it (how poorly we predict the value of ideas, how review capacity fails to keep pace with generation), historical analogies (the spreadsheet era), and organisational research on how AI adoption plays out at scale (DORA). Taken together they make the argument plausible rather than proven, and I would rather say that plainly than let a pile of statistics imply more than it shows.
2. The economics: from scarce engineering to scarce attention
It is worth being honest about how much filtering the old economics was doing on our behalf. Ronny Kohavi, who built the experimentation platforms at Microsoft, Amazon and Airbnb, has reported that when ideas at Microsoft were evaluated in properly controlled experiments, only about a third improved the metric they were designed to improve, with the remainder either flat or negative (Kohavi et al.). At Airbnb his team tested 250 search ideas and found that only 20 had a positive impact on key metrics (AB Tasty interview). The point is not that people have bad ideas; it is that intuition is a poor predictor of value, and the cost of building acted as a crude, expensive but real first pass at separating signal from noise.
AI removes that first pass, and in doing so it displaces the scarcity rather than eliminating it. Previously organisations allocated scarce engineering capacity. Increasingly they will have to allocate scarce attention, integration capacity, operational ownership and organisational commitment, none of which AI has made noticeably cheaper. When the marginal cost of building approaches zero, the cost of deciding what not to build becomes the cost that matters, and the cost of reconciling alternatives has not fallen just because the cost of generating them has.
flowchart LR
subgraph before["Before AI"]
direction LR
A1["Many ideas"] --> B1["Expensive execution<br/>acts as a natural filter"] --> C1["Few solutions"] --> D1["Scarce resource:<br/>engineering capacity"]
end
subgraph after["With AI"]
direction LR
A2["Many ideas"] --> B2["Cheap execution<br/>weak natural filter"] --> C2["Many solutions"] --> D2["Scarce resource:<br/>attention, ownership<br/>and decisions"]
endAn illustrative model rather than measured quantities: the constraint moves from building to deciding.
3. The laptop product portfolio
Imagine walking into a large organisation six months from now. Across the business there are hundreds of small applications: some generated during workshops, some built to demonstrate ideas to leadership, a few solving genuine operational frustrations. Many look polished, with convincing interactions and just enough working logic to feel real. They live on individual laptops, in personal repositories, on temporary hosting and in cloud accounts nobody in technology knows about. There are three versions of the same customer journey, five competing approaches to the same operational problem and a dozen dashboards built on slightly different definitions of the same metric. Some run on synthetic data, others on manually uploaded spreadsheets, and a few have quietly become part of somebody’s daily workflow without anyone deciding that they should.
The direction of travel here is not speculative, even if the magnitude is. Gartner has reported that 41% of employees acquired, modified or created technology outside IT’s visibility in 2022, and predicted that this would rise to 75% by 2027 (Gartner, via Valence Security). That forecast predates mainstream AI coding tools and was driven mostly by SaaS adoption, so it says nothing directly about AI prototypes; it simply shows that technology was already escaping central visibility before generation became cheap.
We have also run a version of this experiment before. Spreadsheets were the original democratisation of building, and Raymond Panko at the University of Hawaii compiled field audits by firms including Coopers & Lybrand and KPMG in which 49 of 54 operational spreadsheets audited between 1997 and 2000 were found to contain significant errors (The Register, summarising Panko). That is a small set of spreadsheets selected for audit rather than a representative sample, so it does not mean nine in ten spreadsheets are wrong; what it shows is how often errors surface once someone actually looks. Panko’s view was that spreadsheet error rates look much like error rates in conventional programming, which is exactly the point: people were doing real software development without the testing, review and ownership that conventional development assumes. AI generated applications have the potential to be that problem again, except with network access, integrations and far more convincing user interfaces.
What makes this portfolio expensive is not storage, compute or even the time it took to generate each item. The real cost is attention: the time other people spend reviewing a prototype, comparing it with alternatives, reconciling its assumptions with everyone else’s and deciding what happens next, together with the steady possibility that someone will misunderstand its status, depend on its behaviour or assume it represents an agreed direction. You might expect cheap prototypes to be easy to delete, since nobody is emotionally invested in something that took an hour, but low attachment does not produce cleanup on its own. When something costs almost nothing to keep, there is no forcing function to decide about it, so it lingers. A prototype on somebody’s laptop can be a perfectly good experiment, but until its purpose, owner, status and the decisions it informed are visible, the organisation has accumulated activity without accumulating knowledge.
4. A prototype is not a product, and usually should not become one
A generated application can look production ready long before it is operationally ready, with good screens, realistic interactions and even working integrations, while every difficult question remains open. Who owns it, what happens when it fails, whether the data is trustworthy, whether it meets security and privacy requirements, who supports it at two in the morning and whether anyone actually needs it are not implementation details to tidy up later; they are what makes software a product rather than a demonstration. AI has made the visible portion of software dramatically easier to produce while the invisible portion remains as hard as ever, and our sense of progress is not a reliable instrument for telling the two apart. METR’s randomised trial with experienced open source developers is a useful warning here: they believed AI had made them around 20% faster when tasks had in fact taken 19% longer (METR). That study is about repository tasks, not prototypes, but it shows how large the gap between felt and measured productivity can be.
None of this is an argument against prototyping, and overcorrecting with approval gates would throw away one of the most valuable things AI gives us. Written requirements are poor at communicating intent: a paragraph describing a customer experience leaves enormous room for interpretation, and stakeholders quietly imagine different journeys and edge cases that they only discover when engineering delivers something none of them pictured. A working prototype makes those assumptions visible early, becomes a shared language between business, design, engineering, risk and operations, and can compress weeks of abstract discussion into a short and concrete conversation.
There is good research supporting the idea that several cheap prototypes beat one polished one. Steven Dow and colleagues at Stanford ran a controlled study in which participants designed web banner adverts either serially, getting feedback on one design at a time, or in parallel, creating several designs before receiving feedback on all of them (Dow et al., ACM TOCHI 2010). The parallel group produced adverts that performed better on click through rates and expert ratings, independent raters judged their designs more diverse, and participants reported a larger gain in confidence in their own design ability. It was a banner advert task rather than enterprise software, so I hold it loosely, but the mechanism transfers well: comparing several alternatives shifts critique away from defending one idea and towards learning about the problem. AI makes parallel prototyping almost free, provided the parallel paths are brought back together.
The distinction I would make most prominent is that a prototype is not necessarily an early version of the eventual product. The phrase “prototype to production” implies that the generated code is expected to evolve into the production system, and in most enterprise settings, and certainly in a regulated one like banking, the right outcome is to preserve the behaviour and discard the implementation. It helps to name three distinct states, because the trouble starts when one is mistaken for another.
| Experiment | Visual specification | Production implementation | |
|---|---|---|---|
| Purpose | Answer a question | Align people on intended behaviour | Deliver sustained value |
| Data | Synthetic or controlled | Synthetic or controlled | Governed production data |
| Architecture | Disposable | Disposable, deliberately never promoted | Maintainable, secure and observable |
| Owner | Whoever asked the question | The team that will deliver the product | An accountable product owner |
| Success means | Learning | Shared understanding and a decision | Measurable adoption and value |
| End state | Retire, keeping the insight | Retire once the real build replaces it | Operate, evolve or decommission |
Confusing these states causes damage in every direction: an experiment survives indefinitely because nobody decided to retire it, a visual specification is mistaken for an implementation and quietly put in front of real users, or every promising prototype is treated as a production candidate and engineering effort is wasted hardening things that should have been thrown away. The solution is not to make every prototype more robust, but to make its state explicit from the moment it is created.
5. Convergence is the new bottleneck
Historically the bottleneck was execution. AI is loosening that constraint, and the evidence suggests the bottleneck does not disappear so much as move downstream, into the places where humans have to review, reconcile and decide. Faros AI analysed telemetry from more than 10,000 developers across 1,255 teams and found that developers on teams with high AI adoption completed 21% more tasks and merged 98% more pull requests, while pull request review time rose by 91% (Faros AI). It is a vendor study and observational rather than experimental, so I treat the exact figures with care, but the shape will be familiar to anyone who has run an engineering organisation: generation scaled, and the human capacity to evaluate what was generated did not.
Now scale that pattern from pull requests to whole solutions. Imagine five teams independently prototyping a new customer onboarding journey. Each produces a polished, clickable flow; each makes different assumptions about which identity checks run upfront, what the customer is asked for and when, and how exceptions are handled; and each wins the support of its own stakeholders, who have seen it work. Six months later the organisation does not have one better onboarding journey. It has five competing interpretations of what onboarding should be, each with a sponsor, and nobody has decided which becomes the standard. Every one of those teams was productive, and the organisation is further from a decision than when it started.
The important distinction is between healthy divergence and unhealthy fragmentation, because from the outside they look identical. Five prototypes exploring five different hypotheses about onboarding are valuable; that is exactly the parallel exploration Dow’s research supports. Five independently maintained solutions to the same problem are not. The difference is not the number of prototypes but whether there is a deliberate mechanism to compare them, learn from them, decide between them and retire the rest.
flowchart LR
subgraph frag["Unhealthy fragmentation"]
direction LR
P1["Onboarding<br/>problem"] --> F1["Prototype A"] --> R1["Quietly in use"]
P1 --> F2["Prototype B"] --> R2["Quietly in use"]
P1 --> F3["Prototype C"] --> R3["Quietly in use"]
end
subgraph conv["Healthy divergence"]
direction LR
P2["Onboarding<br/>problem"] --> H1["Hypothesis A"] --> X{"Convergence owner<br/>compares the evidence"}
P2 --> H2["Hypothesis B"] --> X
P2 --> H3["Hypothesis C"] --> X
X --> OUT["One agreed journey<br/>plus retired alternatives"]
endThat leads to the idea I would put at the centre of any response, which I think of as a convergence obligation. When two or more prototypes address the same problem, someone must be accountable for reconciling what they learned and deciding what happens next. That person should be named rather than implied, and in practice it is usually whoever owns the customer journey or business domain in question, not the builders of any individual prototype. The obligation should be triggered by something concrete, such as a second prototype being registered against the same problem or a prototype reaching its expiry date. And the decision should rest on evidence rather than on whose demo was most impressive: the question each prototype set out to answer, what it actually learned, what data it relied on and who, if anyone, has started to depend on it. The failure mode is not having many prototypes. It is allowing divergence to become permanent because convergence was nobody’s job.
When that happens, the organisation accumulates what I would call convergence debt: the unresolved decisions, competing assumptions and duplicated solutions left behind by unconstrained prototyping. It behaves a lot like technical debt, in that it is cheap to take on, invisible on any single day and expensive to pay down later, but it is debt in collective understanding rather than in code. Five onboarding prototypes are not just five codebases to maintain; they are five different beliefs about what onboarding should be, held by five groups of people, and every month they go unreconciled makes the eventual conversation harder.
6. An operating model for convergence
What does a healthier approach look like in practice? Not a central committee approving every experiment, and not a heavy governance process that destroys the speed AI has unlocked, but a lightweight lifecycle that everyone can name the stage of, with a deliberately broad set of outcomes at the end. The key change from a conventional pipeline is that “Decide” does not mean a binary choice between shipping and failing; it means assigning one of four explicit dispositions, three of which involve no new production code at all.
flowchart LR
E["Explore<br/>test alternative hypotheses"] --> D["Demonstrate<br/>make assumptions visible"]
D --> C["Compare<br/>consolidate the learning"]
C --> X{"Decide<br/>assign a disposition"}
X --> P["Productise<br/>fund and engineer properly"]
X --> S["Specify<br/>keep the behaviour,<br/>discard the code"]
X --> M["Combine<br/>merge the best learning"]
X --> R["Retire<br/>archive the insight,<br/>delete the code"]
M --> C
S -.->|"informs the build"| P
E -.->|"expiry date reached,<br/>no decision"| RRetirement in this model is a normal and successful outcome, not the undesirable branch of a binary decision. A prototype that clarified a requirement and was then deleted has done its job completely, and a culture that only celebrates things that ship will quietly pressure people to ship things that should not. A handful of practices make the lifecycle work.
Govern exposure and dependency, not experimentation itself. This is the principle I would put above all the others. A prototype running locally on synthetic data is not an organisational liability, however many of them exist; a prototype connected to customer data, production APIs or an operational process is a different matter entirely, whoever built it and however small it is. So anyone should be able to explore an idea without asking permission, but within a safe sandbox: synthetic or approved data, approved tools and no uncontrolled integrations with production systems. Visibility should be lightweight, but the moment a prototype touches real customer data, a production integration or real users, that is a deliberate decision with an accountable owner. Without this qualification, “experiment freely” becomes an endorsement of exactly the shadow IT problem described above.
Visible by default, with an expiry date. Every prototype intended to inform an organisational decision should be registered in a shared catalogue with a named owner, a problem statement, the question it is answering and its state. The catalogue is not there to control creativity; it is there so the fifth onboarding team can see the other four before they begin, and so the convergence obligation can trigger. I would also give every prototype a half life when it is registered, a date by which it must be productised, kept as a specification, combined or retired. When the date arrives without a decision, the default should be retirement rather than survival. The lesson from both shadow IT and spreadsheets is that visibility has to be cheap or people will route around it.
Every prototype answers a question. Can this journey be simpler? Will customers understand this interaction? Can this process be automated? Kohavi’s data is the useful reminder here: if most ideas do not move the metrics they were designed to move, the value of a prototype lies in how quickly it tells you which kind of idea you have. If nobody can articulate what a prototype is meant to teach, it is probably being built for the satisfaction of building.
Productise means a handover, not a promotion. Moving from Decide to Productise should mean a deliberate handover into a proper engineering path with the right architecture, controls and support, informed by the prototype’s behaviour, not a prototype being promoted in place because it happened to be running.
Measure convergence quality, not just speed. Counting prototypes, lines of generated code or applications built tells us very little, and the Faros and METR findings show how easily activity metrics rise while delivered value does not. The measure I would lead with is decision throughput, meaning the time between an idea being demonstrated and a decision being made about it, but on its own that rewards premature decisions, so I would pair it with two others: the proportion of prototypes that reach an explicit disposition rather than simply fading, and the proportion of prototypes in operational use that have a named owner. Together those describe how well an organisation converges, not just how quickly.
7. Build more to learn, run less
The DORA research programme at Google offers a useful lens for why this matters at the level of the whole organisation. Its 2024 report, based on more than 39,000 respondents, estimated that a 25% increase in AI adoption was associated with a 1.5% decrease in delivery throughput and a 7.2% decrease in delivery stability, even as individuals reported productivity gains (Google Cloud, 2024 DORA report). The 2025 report, based on nearly 5,000 respondents, found that the relationship with throughput had turned positive while the negative relationship with stability remained, and concluded that AI acts as an amplifier, magnifying the strengths of high performing organisations and the dysfunctions of struggling ones (Google Cloud, 2025 DORA report). These are survey based associations rather than causal proof, but the amplifier framing fits prototype bloat precisely: an organisation with clear ownership, good platforms and a habit of making decisions will turn cheap prototypes into faster learning, while an organisation without those things will turn them into more of what it already struggles with.
So the productivity question for an AI enabled organisation is no longer simply how much software its people can produce. It is how quickly the organisation can turn exploration into shared understanding, shared understanding into decisions and decisions into outcomes. We should encourage more people to prototype, more alternatives to be explored and more assumptions to be challenged, while also expecting most prototypes to disappear once they have served their purpose. The goal is not to build fewer things because building is expensive; it is to run fewer things because permanence should be earned. That leads to a simple principle: prototype broadly, converge deliberately, productise selectively and retire ruthlessly.
The most distinctive conclusion I draw from all this is not that AI will produce a great deal of disposable software, which is increasingly obvious. It is that the organisations best positioned to benefit from AI may be the ones that become exceptionally good at discarding software. AI has made it extraordinarily easy to build another solution, and leadership now has to make it equally normal to decide that we do not need one.
References
- Google Cloud, “Highlights from the 10th DORA report” (2024 Accelerate State of DevOps Report). https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report
- Google Cloud, “Announcing the 2025 DORA Report: State of AI-assisted Software Development”. https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
- METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (July 2025). https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Faros AI, “The AI Productivity Paradox Report 2025”. https://www.faros.ai/blog/ai-software-engineering
- Gartner, “Gartner Unveils Top Eight Cybersecurity Predictions for 2023-2024” (March 2023), as cited by Valence Security. https://www.valencesecurity.com/resources/blogs/gartner-saas-applications-outside-of-it
- Dow, Glassco, Kass, Schwarz, Schwartz and Klemmer, “Parallel Prototyping Leads to Better Design Results, More Divergence, and Increased Self-Efficacy”, ACM Transactions on Computer-Human Interaction, 2010. https://cs303.stanford.edu/papers/ParallelPrototyping2010-final.pdf
- Kohavi et al., “Controlled Experiments on the Web”, KDD 2009 tutorial (Microsoft experiment outcomes). https://ai.stanford.edu/~ronnyk/2009-06-28KDDTutorialT4part1.pdf
- AB Tasty, “1,000 Experiments Club: A Conversation With Ronny Kohavi”. https://www.abtasty.com/?p=79681
- The Register, “Buggy spreadsheets: Russian roulette for the corporation” (2006), summarising Raymond Panko’s field audit research. https://www.theregister.com/2006/05/03/buggy_spreadsheet