The answer is in the Portfolio
Eight figures on AI pilots, nothing on the P&L. Why “double down, redirect, or stop” is the wrong question.

At some point in the next few months, if it hasn't happened already, your board will ask the question that boards everywhere are now asking: we've spent two years and eight figures on AI pilots, nothing has reached the P&L. Do we double down, redirect, or stop?
It sounds like a hard decision. It isn't. It's an unanswerable one. The fact that it's unanswerable is the most important finding.
The false questions
Double down, redirect, or stop treats the pilot portfolio as one thing, deserving one verdict. It never is. An eight-figure portfolio is three or four distinct populations wearing a single budget line: pilots that failed because the idea was wrong; pilots that worked but were aimed at nothing a CFO could see; pilots that proved value and then fell off a funding cliff nobody had planned to bridge; and (usually the largest group) pilots about which no one can honestly say anything at all.
A single verdict applied to that mixture is wrong for most of it. Double down rewards the failures alongside the winners. Stop kills proven value along with the noise. Redirect assumes you know which is which.
You don't. That's the finding.
Four better questions (in order)
The way through is not a verdict but a triage. Every pilot in the portfolio walks the same four questions, in the same order. If the question can't be answered, that's an opportunity to filter it out.
Question one: can it be assessed at all? Is there a baseline that can or has been measured, and does that resemble production? In most portfolios, a third to half of all pilots fail here. This is often invisible and missed by executive sponsors in a rush to implement. These pilots aren't failures. Their value just can't be measured, because they were never set up to be assessed effectively. Many were never really pilots at all: a pilot with no scale hypothesis and no pre-committed path to production is a demonstration. The organisation wasn't buying evidence. It was buying reassurance.
Question two: did it prove a mechanism of value? Not "did the model work". Did the pilot demonstrate a causal link between the capability and a business effect, including the part everyone forgets to test: did human behaviour actually change? Technical success with no adoption = failure. Failing here is the only genuine stop, and it's the healthy one. A well-instrumented disproven hypothesis is an asset. If the learning is harvested formally, the write-off bought something.
Question three: does the value land on a number someone owns? The mechanism works. But does the benefit arrive on a line item the CFO already reconciles, owned by a named executive? Here is where most of the quiet disappointment lives. Pilots aimed at diffuse improvement — a dozen workflows nudged up a few per cent each — produce value that is real and invisible in the same breath. If the value only shows up in metrics nobody reconciles, at board level it doesn't exist. The test is brutal and takes ten seconds: name the line item and name the person. If you can't, the verdict is redirect — the capability is sound, the aim point isn't. This is usually the cheapest value in the whole portfolio, because the technical risk is already retired. It just needs pointing at something that counts.
Question four: do the economics survive production? Value proven, number owned. Now: does the cost per transaction hold at ten million events, not a thousand? Does the data exist without a team hand-curating it? Does it run inside the real workflow, with an operating owner on the other side? Pilots that fail here are the ones the board most consistently misreads. They look like they need more time. They don't. The pilot is finished — it succeeded. What needs funding is the scaling system around it: the platform, the process redesign, the change effort, typically several multiples of what the pilot cost. The verdict is not "extend the pilot." It is fund the cliff — the cliff that was always going to be there, that nobody budgeted for, because the pilot was approved as an experiment rather than as the first stage of an investment.
What survives all four questions scales now. In a typical portfolio, it is a small handful. That rarity is not an indictment. It is a prioritisation gift: the survivors have earned concentration of capital, not an equal share alongside everything else.
My experience tells me that we're likely to skip question one in a rush to demonstrate value. Two we manage to fit a definition of value, while three can work, but is often rejected if the answer to two isn't real. Question four is even more challenging, it's not just about technology and if the measurement wasn't setup to include "real" adoption, it can fail before it even starts.
Read the distribution, not the pilots
Here is where triage becomes strategy. Once every pilot has a verdict, look up from the pilots and read the shape of the portfolio, because the distribution diagnoses the system that produced it.
A portfolio that is mostly un-assessable has a measurement discipline problem: fix the pilot template before approving another dollar. Mostly redirect: a targeting problem. Use cases are being selected for visibility or feasibility rather than for owned numbers. The evidence on this is uncomfortable; budgets have flowed disproportionately to the most visible functions, where returns have been thinnest, while the unglamorous back office quietly pays back. Mostly stuck at the cliff: a funding architecture problem. Pilots were never designed as staged investment decisions with scale capital reserved against success. Mostly stopped at gate two: the problem is upstream, in how ideas reach the pipeline at all.
This is the part the board actually needs. The pilot verdicts recover this portfolio. The distribution reading prevents the next eight figures producing the same shape.
About that famous number
You will have seen the statistic. MIT's finding that ninety-five per cent of enterprise AI pilots deliver no measurable P&L return. It is quoted everywhere, usually as proof that AI disappoints. Its methodology has been fairly criticised: a narrow definition of success, a short window, blindness to value that lands anywhere other than the P&L. Both things are true, and together they make the real point better than either does alone. The number is less a measurement of AI's failure than of enterprises' inability to measure. The scandal is not that the pilots returned nothing. It is that, for most of them, nobody can say. It should never have passed question one.
So when the board asks, double down, redirect, or stop? The honest answer is: that's a verdict, and what you need first is a sort. Triage the portfolio. Fund the cliffs you find. Re-aim the mistargeted. Stop what was disproven and keep the learning. And treat the distribution as the real report: it will tell you, more precisely than any strategy review, why your organisation converts capital into pilots and pilots into nothing. And more importantly, what to change so the next tranche converts into earnings instead.
The pilots were never the problem. The system that produced them is the thing under review.
If your board is circling double down, redirect, or stop, that's the wrong question. LOWEMGMT helps executive teams triage the portfolio, fund the cliffs, and re-aim what's mistargeted before the next tranche repeats the same shape. Talk to us about your portfolio.
