Every AI governance document says "human in the loop". Very few say where, and the default answer — everywhere — produces a control that stops working on a predictable schedule.
How approval queues fail
The mechanism is not mysterious.
A queue is created. Reviewers give each item real attention because the volume is low and the system is new. Volume grows. Attention per item falls. At some point the median review time drops below the time required to read the case, and from then on the queue is a rubber stamp.
The control still exists on the org chart. It still produces an audit trail. The audit trail is now evidence that oversight happened when it did not, which is strictly worse than having no control at all, because it is load-bearing in somebody's risk assessment.
We have watched this happen in about ten weeks, repeatedly, across very different organisations. It is not a culture problem.
Where oversight actually belongs
At the policy boundary. Cases just inside or outside a threshold are where automated systems make their most expensive errors, and where human judgement adds most. Route them by construction.
Where evidence lanes disagree. If retrieval supports one conclusion and the deterministic checks support another, that disagreement is the signal. Do not let a confidence score average it away.
Where a party is vulnerable. Not a confidence question at all. A policy question, decided in advance, enforced in the runtime.
On a blind random sample. A fixed percentage of autonomous decisions, re-worked by a human who cannot see what the system decided. Disagreements go to a governance group.
That last one is the control that tells you whether the others are working, and it is the one most often missing.
Why blind samples work when queues do not
A queue reviewer sees the system's answer first. Anchoring does the rest — the question stops being "what is right" and becomes "is this obviously wrong", which is a much weaker test.
A blind reviewer produces an independent judgement. Comparing the two gives you a real measure of agreement, and the disagreements are genuinely informative rather than being a list of things someone did not bother to challenge.
Sample rates of 1–5% are typical. Resist reducing the rate when agreement is high — that is exactly when the measurement is cheap and when drift will be hardest to notice.
Make the reviewer's job possible
Where humans do review, the case must arrive assembled: the evidence, the lineage, comparable prior decisions, what the system concluded and why, and what happens next under each option.
If reviewing means opening four other systems, the review will be shallow regardless of intent. Time a handful of your own reviews and split them into retrieval and judgement — in most operations the first number dwarfs the second, and it is the one you can remove.
Give reviewers somewhere to put disagreement
If a reviewer thinks the system is wrong in a way that will recur, there must be a route for that observation that is not a Slack message.
Overrides should be structured, categorised, and fed into the evaluation suite as candidate cases. A reviewer who sees their objection become a test case stays engaged. One who does not, stops objecting.
The goal is not maximum human involvement. It is oversight that is still real in month eighteen. Those are different designs, and only one of them survives volume.
- Governance
- Agents
- Human Oversight