COMIT · AI Track · September 2026

Why your
AI pilot
will fail

Six real pilots. Not one of them failed on the model. The seven questions that would have caught all six, and every number in the talk with its source.

The readiness test

Seven questions to ask before you spend a penny

Every one of these is answerable before you fund anything, and every one of them comes from a pilot that failed because nobody asked it.

01

Whose hours are we removing, and how many?

Name the person and the number. “Efficiency” is not a number, and a business case that cannot name an hour is not a business case.

From: the pilot that tried to replace the whole team
02

Who checks the output, and how long does that take?

Subtract the checking from the saving. If nothing is left, you have not found a pilot yet. A tool that is right most of the time, with no way of telling which times, buys you an expensive second opinion.

From: the pilot that saved nobody any time
03

What must it cost, and how fast must it run?

Pounds per document. Seconds per record. And where the data is allowed to sit. Write them down before the pilot rather than after it, then put telemetry on latency and cost per run, or you will not find out until you are live.

From: the model that was good, and too slow
04

Does this actually need AI?

The cheapest question here and almost nobody asks it. Plain code is faster, costs less and gives you the same answer twice. The answer is nearly always a hybrid.

From: the pilot that did not need AI
05

Who owns this, by name?

Not a department. A person whose job it is, with someone senior standing behind them. A pilot with no owner is not a pilot, it is a hobby.

From: the pilot nobody was told about
06

What are people already using to do this today?

Ask before you build. Half your pilot may already exist, in the dark, on somebody's personal account. Run an amnesty and treat it as free research rather than a policy breach.

From: the pilot everyone had already run without you
07

What result at week eight makes us stop?

The one everybody skips

Agree how it dies before you agree what it costs, while everyone still likes each other. Nobody can write an honest kill criterion once the money is spent and somebody's reputation is attached to the answer. Decide the number now, write it down, and give the person who owns the pilot permission to hold you to it.

If you remember one thing

Not one of them failed on the model.

Every one of those systems worked, at least somewhere. They failed on the business case, on a constraint nobody had written down, and on people who were already ahead of us. Every one of those is checkable before you spend anything, and not one of them needs a better model to fix.

Part one

The six pilots

Real engagements. Some we built, some we were called in to rescue. No client is named and none will be.

01

The pilot that tried to replace the whole team

The scope was wrong. It automated the judgement and left the humans the legwork. Do it the other way round. The cheapest win in AI is usually deleting input rather than buying a bigger model.

02

The pilot that saved nobody any time

Excellent on a tuned test set, which it would be, because it had been tuned until it was. Verifying it in the real world cost exactly what producing it saved.

03

The model that was good, and too slow

Good enough on quality. Too slow and too expensive to run for real. A far less impressive model on ordinary hardware won the job, because it was fast enough. Accuracy is one requirement out of six.

04

The pilot that did not need AI

A sorting problem with “AI” written on the funding request. Ordinary fuzzy matching did the heavy reduction; only the genuinely ambiguous cases ever reached a model.

05

The pilot nobody was told about

Well built, well demoed, never opened. Every technical box ticked and every human one missed: nobody was trained, nobody was told, and nobody's name was against it.

06

The pilot everyone had already run without you

Your pilot is not competing with the status quo. It is competing with the tab already open on somebody's second monitor, doing a version of the job today, free. Make yours better than the free tool, or people go round the side of it.

Provenance

Every number in the talk

13% · 35% · 58% — AI adoption by sector

ONS, Artificial intelligence in UK businesses: 2023 to 2026

Construction 13%. All UK business 35%. Information and communication 58%. Adoption across all UK business has risen from around 12% to around 35% since 2023.

Business Insights and Conditions Survey wave 159, released 20 July 2026. Fieldwork 5 to 28 June 2026, 38,637 businesses responding, response rate 26.7%.
ons.gov.uk
Caveat: “use of at least one AI technology” is a broad definition, and the survey covers businesses with ten or more employees only. Do not blend this with DSIT's figures, which use different fieldwork dates and question wording; cite one or the other, never an average.
39% pilot · 19% against 29% in regular use — the conversion gap

RICS, AI in Commercial Property and Construction Report 2026

Construction and commercial property both run early-stage pilots at 39%. Only 19% of construction respondents reach regular use in a specific process, against 29% in commercial property. RICS attribute the gap to construction data resetting at every project boundary, where commercial property runs on portfolio-level recurring data.

Published 18 August 2026. n=3,148 (Global Construction Monitor 1,883; Global Commercial Property Monitor 1,265), fieldwork Q1 2026.
rics.org
Caveat: respondents are RICS members, who are likely to be more digitally engaged than the sector average, so the true conversion rate is probably worse rather than better.
78% — people bring their own AI tools to work

Microsoft and LinkedIn, 2024 Work Trend Index Annual Report

78% of people who use AI at work bring their own tools to do it. Alongside it: 75% of knowledge workers use AI at work, and 46% of those users started less than six months before the fieldwork.

Published 8 May 2024. Approximately 31,000 respondents across 31 countries, plus LinkedIn labour and hiring data and Microsoft 365 telemetry.
microsoft.com
Two caveats: the denominator is AI users rather than all workers, so the correct phrasing is “of people who use AI at work”. And it is May 2024 fieldwork. The direction of travel since probably makes it conservative, but the date matters.
50 million comparisons

Arithmetic, not a survey

Comparing 10,000 records with each other is n(n−1)/2, which is 49,995,000 pairs. At any price per model call, that is not a pilot. It is a funding round.

And the ones we left out

The numbers I did not use

These get quoted at every AI conference in the country. All of them are weaker than they look, and you should know why before you put one in a board paper.

The famous 95%

MIT NANDA, The GenAI Divide: State of AI in Business 2025

The report's own headline is 95% of organisations are getting zero return, and not 95% of pilots. It is labelled preliminary, it is not peer-reviewed, and it never fully reconciles how 95% is derived from its underlying data: a review of 300-odd public initiatives, 52 structured interviews and 153 survey responses collected at four industry conferences.

The qualitative message about integration and workflow failure holds up well. The precise number does not bear weight.
Gartner's 30%

Gartner press release, 29 July 2024

“At least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025.” This is a forecast about 2025 issued in 2024, rather than a measured outcome, and nobody has published a retrospective confirming whether it happened.

Word it as a prediction, or leave it alone.
RAND's 80%

RAND, The Root Causes of Failure for Artificial Intelligence Projects (RRA2680-1, 2024)

RAND's own wording is “by some estimates, more than 80% of AI projects fail”. That is RAND repeating an external estimate as framing, rather than a figure RAND measured. Their actual fieldwork was 65 interviews with data scientists and ML engineers, producing five root causes of failure.

Cite RAND for the causes. Never for the 80%.
One more, while we are here

“80% of a data scientist's time goes on data preparation”

Folklore. It traces to a 2014 expert estimate quoted in a newspaper and a 2016 survey in which 60% named cleaning as the most time-consuming task and a separate 19% named collection. The 80% is press arithmetic. A later Anaconda survey put preparation nearer 45%.

The offer from the stage

Bring me your stuck pilot.

Thirty minutes, no charge. Your pilot, these seven questions, and my honest answer about which of the six it is.

We build AI and data systems into construction, infrastructure and property businesses, and our engineers work inside your systems for as long as that is useful and no longer. We will tell you when you do not need us.

Or tell me about it here

Alex reads these himself.