Your AI pilot failed before anyone opened a laptop

Many AI pilots produce no measurable benefit. The model is rarely the whole problem: the tool was bought before anyone looked at what was broken in the work.

Every company now has an AI pilot, and most will soon have an AI pilot post-mortem. An MIT research group claimed last year that 95 percent of enterprise generative AI pilots deliver no measurable impact on the bottom line.1 The figure has been criticised on fair grounds, and I will get to that. The direction, though, is familiar to anyone who has watched a licence get bought, a demo get applauded, and the actual work continue exactly as before. Sometimes a pilot dies on the technology itself; the model simply cannot yet do what the salesperson promised. Far more often the pilot was dead at the moment of purchase, because nobody had defined the problem it was supposed to solve. The tool arrived before the question.

A pilot is a brief that skipped the diagnosis

An AI pilot is usually born the same way a bad brief is. Somebody reads a competitor’s press release, the board asks about an AI strategy, and a project appears in the calendar with a name but no problem. The purchased solution is a guess, made before anyone looked at the work it was meant to change.

The pattern is older than the technology. A rebrand gets ordered when sales are flat. An app gets commissioned when support drowns. The name of the deliverable exists before the diagnosis does, and the name points the wrong way. AI changed nothing about this mechanism; it just gave it a newer and more expensive vocabulary.

What the 95 percent actually says

Honesty first: the MIT report rests on a thin sample, its definitions are loose, and it is not peer-reviewed research.2 The precise number should not be carried into any meeting. The report’s most useful finding is not the percentage but the cause it points at: the projects that worked started from one real workflow, while the ones that failed started from a tool that went looking for a use afterwards.

The same split shows up here in Finland, where we work. AI use is growing fast even in small firms, but the use is shallow, and understanding of things like regulation trails far behind.3 Plenty of companies can open a chat window. Few can name the piece of work it changed, or show the change in a number.

The question the pilot never asked

A useful pilot starts from a question that does not mention AI at all: what currently makes your work harder than it should be. The answer is almost always concrete. A quote is assembled by hand from three files. The same customer question gets answered by email four times a week. Invoicing slips because the hours live in someone’s memory.

Some of these are excellent AI problems. Others are solved by a form, a price list, or one phone call, and in those cases buying AI would have been the most expensive available way to avoid the cheapest fix. Diagnosis is not a delay. It is the only way to know which group your problem belongs to.

Small, measurable, and ugly

Once the problem has a name, a good pilot is almost always a disappointment as presentation material. One workflow, one metric, a few weeks. The baseline gets measured before the start, because without it every pilot is declared a success afterwards and nothing changes.

The dullness is a feature. A project that fits in one sentence is also a project whose failure is visible. Three months in, you know whether the hour spent building each quote came back or not. That is more than most strategy decks ever tell you.

The inconvenient part

Hype punishes the buy-first-diagnose-later company twice: first in money, then in the cynicism that blocks the next attempt, which might have been the right one. The third cost is the quietest. A failed pilot teaches an organisation to say we tried that and it didn’t work, when nothing was really tried; a tool was tried, without a question attached. Next year’s model will be better again, and it will still be for sale then. Your own work, and the friction inside it, is available to read today, for free. Which one will you look at first?

From words to done.

Sources

  1. MIT NANDA, “The GenAI Divide: State of AI in Business 2025”; summary via Fortune. https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html
  2. Marketing AI Institute, critique of the report’s sample and definitions. https://www.marketingaiinstitute.com/blog/mit-study-ai-pilots
  3. AI Finland, “AI in Finnish Business 2026”, and the FAIR survey on regulatory readiness. https://aifinland.fi/en/ai-in-finnish-business-2026/ and https://www.fairedih.fi/en/2026/04/30/finnish-firms-embrace-ai-rapidly-but-regulatory-knowledge-lags-far-behind-survey-finds/

Sound familiar? A 25-minute conversation commits you to nothing.

Book a time