Six real pilots. Not one of them failed on the model. The seven questions that would have caught all six, and every number in the talk with its source.
Every one of these is answerable before you fund anything, and every one of them comes from a pilot that failed because nobody asked it.
Name the person and the number. “Efficiency” is not a number, and a business case that cannot name an hour is not a business case.
Subtract the checking from the saving. If nothing is left, you have not found a pilot yet. A tool that is right most of the time, with no way of telling which times, buys you an expensive second opinion.
Pounds per document. Seconds per record. And where the data is allowed to sit. Write them down before the pilot rather than after it, then put telemetry on latency and cost per run, or you will not find out until you are live.
The cheapest question here and almost nobody asks it. Plain code is faster, costs less and gives you the same answer twice. The answer is nearly always a hybrid.
Not a department. A person whose job it is, with someone senior standing behind them. A pilot with no owner is not a pilot, it is a hobby.
Ask before you build. Half your pilot may already exist, in the dark, on somebody's personal account. Run an amnesty and treat it as free research rather than a policy breach.
Agree how it dies before you agree what it costs, while everyone still likes each other. Nobody can write an honest kill criterion once the money is spent and somebody's reputation is attached to the answer. Decide the number now, write it down, and give the person who owns the pilot permission to hold you to it.
Every one of those systems worked, at least somewhere. They failed on the business case, on a constraint nobody had written down, and on people who were already ahead of us. Every one of those is checkable before you spend anything, and not one of them needs a better model to fix.
Real engagements. Some we built, some we were called in to rescue. No client is named and none will be.
The scope was wrong. It automated the judgement and left the humans the legwork. Do it the other way round. The cheapest win in AI is usually deleting input rather than buying a bigger model.
Excellent on a tuned test set, which it would be, because it had been tuned until it was. Verifying it in the real world cost exactly what producing it saved.
Good enough on quality. Too slow and too expensive to run for real. A far less impressive model on ordinary hardware won the job, because it was fast enough. Accuracy is one requirement out of six.
A sorting problem with “AI” written on the funding request. Ordinary fuzzy matching did the heavy reduction; only the genuinely ambiguous cases ever reached a model.
Well built, well demoed, never opened. Every technical box ticked and every human one missed: nobody was trained, nobody was told, and nobody's name was against it.
Your pilot is not competing with the status quo. It is competing with the tab already open on somebody's second monitor, doing a version of the job today, free. Make yours better than the free tool, or people go round the side of it.
Construction 13%. All UK business 35%. Information and communication 58%. Adoption across all UK business has risen from around 12% to around 35% since 2023.
Construction and commercial property both run early-stage pilots at 39%. Only 19% of construction respondents reach regular use in a specific process, against 29% in commercial property. RICS attribute the gap to construction data resetting at every project boundary, where commercial property runs on portfolio-level recurring data.
78% of people who use AI at work bring their own tools to do it. Alongside it: 75% of knowledge workers use AI at work, and 46% of those users started less than six months before the fieldwork.
Comparing 10,000 records with each other is n(n−1)/2, which is 49,995,000 pairs. At any price per model call, that is not a pilot. It is a funding round.
These get quoted at every AI conference in the country. All of them are weaker than they look, and you should know why before you put one in a board paper.
The report's own headline is 95% of organisations are getting zero return, and not 95% of pilots. It is labelled preliminary, it is not peer-reviewed, and it never fully reconciles how 95% is derived from its underlying data: a review of 300-odd public initiatives, 52 structured interviews and 153 survey responses collected at four industry conferences.
“At least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025.” This is a forecast about 2025 issued in 2024, rather than a measured outcome, and nobody has published a retrospective confirming whether it happened.
RAND's own wording is “by some estimates, more than 80% of AI projects fail”. That is RAND repeating an external estimate as framing, rather than a figure RAND measured. Their actual fieldwork was 65 interviews with data scientists and ML engineers, producing five root causes of failure.
Folklore. It traces to a 2014 expert estimate quoted in a newspaper and a 2016 survey in which 60% named cleaning as the most time-consuming task and a separate 19% named collection. The 80% is press arithmetic. A later Anaconda survey put preparation nearer 45%.
Thirty minutes, no charge. Your pilot, these seven questions, and my honest answer about which of the six it is.
We build AI and data systems into construction, infrastructure and property businesses, and our engineers work inside your systems for as long as that is useful and no longer. We will tell you when you do not need us.
Alex will come back to you within a working day.