How to measure AI value without misleading the organization
Most AI metrics are true and useless: activity counted as outcome, projection reported as result, a pilot cohort extrapolated to a department. Five parts fix all three.
The problem with AI reporting is rarely dishonesty. It is that the numbers are technically accurate and structurally misleading, and the people producing them cannot see it because each individual figure survives a fact-check.
Three patterns account for most of it.
Activity reported as outcome
"12,000 queries this month" tells you the tool was opened. It does not tell you whether anything was decided faster, or better, or at all. Adoption metrics are worth tracking — a workflow nobody uses cannot create value — but they are leading indicators, not results, and presenting them in the outcome slot invites an investment decision the evidence does not support.
The test: could this number go up while the business result stayed flat? If yes, it is activity.
Projection reported as result
A model of expected savings is a legitimate planning artefact. It becomes a problem the moment it appears in the same typeface as an observed figure, because six months later nobody remembers which was which — including the person who made the slide.
Every forward-looking number needs a visible qualifier at the point of display, not in a footnote: projected, target, controlled pilot. If the qualifier makes the number look weaker, that is the qualifier working correctly.
Pilot extrapolated to population
The participating team volunteered. They had the builder's attention, cleaner-than-average data, and a reason to care. Their result is real and it is not the department's result.
This is the single most expensive error in AI reporting, because the investment that follows is sized against a number that was never true at that scale — and the shortfall does not surface until the money is committed.
The five parts
Every executive-facing figure carries all five. Missing any one of them is what makes a number misleading rather than merely incomplete.
- Value — the figure itself
- Label — what it represents, in the reader's language, not the system's
- Qualifier — scope, period, cohort, baseline, evidence level
- Trend — the comparison point, where one exists
- Status — projected, piloting, validated, or production
Assembled, it reads like this:
Notice how much the qualifier removes. It is not the whole service organization, it is not open-ended, and escalated cases — the hard ones — are outside the measurement. A reader can now decide what to do with it, which is the only thing a metric is for.
Baselines are the part everyone skips
None of the above works without a baseline captured before the change. It is unglamorous, it delays the interesting work by a week, and it is the difference between a result and an anecdote.
If the baseline does not exist and the workflow has already changed, say so plainly rather than reconstructing one. "We do not have a reliable pre-launch baseline for this workflow" is a defensible sentence. A baseline assembled after the fact from memory is not, and it will not survive the first person who checks.
What good reporting costs
Roughly a day per initiative per quarter, and a willingness to publish numbers that are smaller than the ones a vendor would put on the same slide. That is the whole cost.
What it buys is an organization that can tell the difference between a programme that is working and one that is merely busy — which is, eventually, the only question that matters.
Continue reading
Human-in-the-loop is an operating model, not a checkbox
Saying a person reviews the output is not a control. A control specifies who reviews, against what standard, with what authority, and what happens when they disagree.
Why most AI pilots never become operating capabilities
A pilot proves a model can do something. Production proves an organization can rely on it. The gap between the two is not technical, and it is where most programmes stop.

Next step
Your organization does not need more AI experiments. It needs a trustworthy path forward.
Start with a structured working session to identify where AI can create value, what is preventing progress, and which next step is justified by the evidence.
No generic transformation pitch. No required platform purchase. No commitment before the opportunity and constraints are clear.
