Inory AI

Human-in-the-loop is an operating model, not a checkbox

Saying a person reviews the output is not a control. A control specifies who reviews, against what standard, with what authority, and what happens when they disagree.

"Human-in-the-loop" appears in almost every AI governance document we are shown. In most of them it is a single line, and it does no work. It reassures the reader without constraining the system, which is the definition of a control that will not hold.

A control that holds answers four questions. Any design that cannot answer all four has a review step in name only.

Who reviews

Not "the team". A role, and behind it a person who can be asked why a decision was made. If the reviewer is whoever happens to be least busy, the standard drifts to whoever is most tolerant of risk, and it drifts silently.

The reviewer also has to be someone whose job leaves time for it. A review step assigned to a saturated team is a rubber stamp with a timestamp — arguably worse than no review, because it creates a record that suggests scrutiny happened.

Against what standard

Review against what, specifically? "Check it looks right" is not a standard, because two reviewers will apply it differently and neither can be wrong.

Useful standards are written as decision rules a second person can apply and reach the same answer:

WeakHolds
Check the summary is accurateEvery figure in the summary appears in the source document
Make sure the classification is sensibleReject if confidence is below the threshold or the case cites more than one policy
Confirm the reply is appropriateNo commitment about timing, price or liability without a named approver

The right-hand column can be audited. The left-hand column can only be argued about.

With what authority

Can the reviewer stop the workflow, or only annotate it? If the system proceeds regardless, the review is telemetry, not control — worth having, but do not call it a safeguard.

Real authority means the reviewer can reject, and rejection has a defined consequence: the case routes somewhere, someone is accountable for the backlog that creates, and the rejection rate is visible to the person who owns the outcome.

And what happens when they disagree

Disagreement is the interesting case, and it is almost always undesigned. The agent proposes, the reviewer rejects, and then what? If the answer is "the reviewer does it manually and moves on", the system never learns and the failure is invisible in aggregate.

The disagreement is the most valuable signal the workflow produces. It should be captured with the reason, reviewed on a cadence, and used to change either the system or the standard. A workflow that cannot tell you why it was overruled last month cannot be improved except by guesswork.

The uncomfortable implication

Designing review properly costs more than building the agent. It touches role definitions, capacity planning, quality standards and management cadence — organizational work, owned by managers, that no model release will do for you.

That is the actual reason human-in-the-loop stays a checkbox. Not because teams do not understand it, but because the version that works is a change to how people are managed, and that is a harder project to sponsor than a change to how software is deployed.

It is also the one that decides whether the system can be trusted with anything that matters.

Continue reading

All insights

Next step

Your organization does not need more AI experiments. It needs a trustworthy path forward.

Start with a structured working session to identify where AI can create value, what is preventing progress, and which next step is justified by the evidence.

No generic transformation pitch. No required platform purchase. No commitment before the opportunity and constraints are clear.