Inory AI

Why most AI pilots never become operating capabilities

A pilot proves a model can do something. Production proves an organization can rely on it. The gap between the two is not technical, and it is where most programmes stop.

A pilot answers one question: can the model do the task at all? That question is usually settled in a fortnight, and the answer is usually yes. Then the initiative stalls — not because anyone decided to stop it, but because nobody wrote down what would have to be true for it to continue.

Production answers a different question, and it is an organizational one: can this workflow be relied on by people who did not build it, on a Tuesday, when the person who did build it is on leave? Everything that makes that answer yes is work the pilot did not do.

What the pilot quietly skipped

A demonstration can borrow a lot. It borrows a clean data extract instead of a live integration. It borrows the builder's judgement instead of a written quality bar. It borrows enthusiasm instead of a named owner. None of that is dishonest — it is the right way to test an idea cheaply. It just means the pilot's success does not transfer.

Look at what has to exist before the same workflow can run unattended:

  • a named business owner who is accountable for the outcome, not for the project
  • a data and system access path that survives a permissions review
  • evaluation criteria that someone other than the author can apply
  • explicit rules for when a person must review, and what they are reviewing for
  • an escalation path for the cases the system should not decide
  • monitoring that surfaces failure before a customer does
  • documentation good enough that the second engineer can change it safely

Two failure modes, and they look identical from outside

The first is the pilot that was never worth scaling. The workflow had low volume, or the benefit was concentrated in one team's frustration rather than in the operating result. This is a success, provided somebody says so out loud and retires it. A pilot that produces a clear no is cheaper than a production system that produces a slow maybe.

The second is the pilot that was worth scaling and was not resourced to get there. The integration work, the review design, the training, the ownership transfer — all of it is real work, and none of it looks like AI work, so it tends to be nobody's budget line.

The two are hard to tell apart because they end the same way: quietly. Distinguishing them requires a baseline that was recorded before the pilot ran, which is the single most commonly skipped step.

Measure the thing you would have to defend

The figure that survives scrutiny is the one that carries its own scope. This is what an honest outcome looks like:

28%Reduction in average handling timeMeasured for the participating service team during the eight-week controlled pilot; excludes escalated cases.

Note what the qualifier is doing. It names the cohort, the period, and what is excluded. That is not hedging — it is the difference between a number an executive can act on and a number they will be embarrassed by in two quarters, when someone asks whether it applied to the whole department.

A pilot result reported as a production result is the most expensive kind of optimism, because the investment that follows is sized against a number that was never true at that scale.

What to do differently on the next one

Before the build starts, write down the decision the evidence has to support. Not the metric — the decision. "If handling time falls by more than X for the participating team over eight weeks, we extend to the other two teams and hire an owner." That sentence forces a baseline, a cohort, a period and a threshold into existence, and it makes the outcome actionable whichever way it lands.

Then resource the unglamorous half. The integration, the review design, the training and the handover are the work that turns a demonstration into a capability. If they are not on the plan, the plan is for a pilot, and it will end like one.

Continue reading

All insights

Next step

Your organization does not need more AI experiments. It needs a trustworthy path forward.

Start with a structured working session to identify where AI can create value, what is preventing progress, and which next step is justified by the evidence.

No generic transformation pitch. No required platform purchase. No commitment before the opportunity and constraints are clear.