A Checkmark Isn’t a Check
Plain-English premise: An approval record proves someone was at the gate. It doesn’t show what they checked, or whether they could have stopped the work.
Judgment Seat: The approver whose name is on the record, whether or not the gate gave them enough to judge.
Episode 1:
👋 Last week’s closing promised an episode called The Button That Used to Be a Person. The research changed my mind. The person didn’t disappear. They clicked. Episode One budgeted the reviewer’s attention. Episode Two asked what the sender owes before review begins. This week is about the moment in between: the approval itself, and what it can and can’t tell anyone who reads it later.
Episode 2:
Research Binder: the receipts (citations + source notes) are compiled in a PDF at the bottom of this post.
⚡ Pressure
AI can expand what arrives for approval, and the click that clears it records presence, not judgment.
🎯 Payoff
This issue gives you the Gate Honesty Check: five questions that test whether an approval gate can support the judgment its record implies.
Pressure Moment
In mid-September, a software engineer posted a short after-action note. He had approved a pull request an AI agent wrote for a system he designed. Tests passed. The diff looked reasonable. He merged it without fully tracing one change. Two days later, a background job quietly started double-processing records. It’s one person’s public account, from software, and it says nothing about how often this happens. What it shows is narrower and more useful. He did check things. The merge record would look identical whether he traced that change or not.
The log recorded a human. It didn’t record the check.
I’ve inherited plenty of “approved” decisions in enterprise environments where the approval was real, but the judgment behind it wasn’t recoverable. We could prove who signed and what moved forward. Reconstructing what they had actually checked sometimes required emails, meeting notes, and conversations with people who were still around. The problem wasn’t necessarily the review. It was that the record couldn’t carry the review forward.
None of this means the approval was empty. Sometimes a green checkmark is the compact end of an excellent review. Sometimes the review was thinner. Sometimes we simply can’t tell, and not knowing is different from knowing that no judgment occurred. An incomplete record is not proof of an empty mind. The useful question is narrower: what evidence justifies treating this particular approval as meaningful oversight?
Start by separating the three questions a single click tends to collapse into one.
Was a human present? Did meaningful judgment occur? Did it improve the outcome? The log can answer the first: who clicked, when, and what state the system recorded. The other two need more evidence: what the person inspected, what evidence they used, and what would have made them refuse.
The one number on that graphic shows why the design around a decision matters. In a 199-participant experiment (Buçinca, Malaya & Gajos, 2021), people using simple explainable AI accepted incorrect advice on one wrong-advice subdecision about 64 percent of the time, versus 48 percent under cognitive-forcing designs. That’s a conditional result, not a 16-point cut in overall failure. It still shows that how a decision is set up changes what people do with bad advice. A human’s presence alone doesn’t guarantee that.
What the Record Can Prove
Organizations like signatures because one field compresses hours of work into “Approved by Taylor Kim.” That compression is useful. It also loses a lot.
A meta-analysis of 106 experiments and 370 effect sizes (Vaccaro, Almaatouq & Malone, 2024) found that human-AI combinations beat humans working alone on average, yet underperformed the better of the two working alone, with results varying widely by task. That finding supports neither “humans make AI better” nor “get humans out of the way.” Performance depends on the task, the system, and how the collaboration is designed. So read the record for what it can support: a sign-off shows presence, a review trail offers observability, and only outcome data speaks to effect.
The gap between a record and what actually happened isn’t new, and it isn’t an AI problem. There’s useful evidence from a setting with no AI at all.
In one hospital audit (Salgado, Barber & Danic, Journal of Patient Safety), independent audio review of 100 surgical cases found lower checklist fidelity than the paper completion records indicated. That doesn’t mean the clinicians were lying, or that no safety work occurred. The audit compared spoken checklist items with paper records, not anyone’s private judgment. Missing detail doesn’t prove missing judgment. It proves the record can’t settle the question on its own.
The answer isn’t more paperwork or surveillance. It’s deciding what each record is for. Approval status tells you who acted. Review evidence tells you what was inspected. Outcome data tells you what happened next. One checkmark can’t do all three jobs.
Watching Isn’t Stopping
Now imagine the reviewer spots a problem. What happens next? Darren Kimura, a technology executive quoted in CIO this month, described many companies that claim a human in the loop as actually having “a human watching the loop”: someone who can flag a concern but can’t stop, change, reject, or escalate it. He’s a vendor voice, so treat the line as a sharp distinction, not a measurement.
The NIST AI Risk Management Framework tells organizations to assign mechanisms to supersede, disengage, or deactivate AI systems that perform inconsistently with their intended use. That’s a governance principle, not proof of outcome benefit. The closest outcome evidence comes from hospitals. A Cochrane review of escalation systems (McGaughey et al.) covered 4 randomized trials with 455,226 participants plus 7 nonrandomized studies. It couldn’t isolate authority as the active ingredient, because pathway design, recognition, communication, response, and implementation were bundled together.
That doesn’t mean authority is irrelevant. It means we should be precise about the pathway. Can the reviewer detect the problem? Can they act on it? Will the system honor the action? An approver without a usable pathway may be an observer with a better title.
Who Decided the Depth?
Episode One dealt with review capacity. This section asks a different question: when more work arrives, who decides how deep each review goes?











