Empathy Engine

Empathy Engine

A Human Clicked Approve

The Log Says a Human Approved It. What Did the Human Check? (JL-003 | Judgment Layer)

Mark S. Carroll's avatar
Mark S. Carroll
Sep 23, 2026
∙ Paid

A Checkmark Isn’t a Check

Plain-English premise: An approval record proves someone was at the gate. It doesn’t show what they checked, or whether they could have stopped the work.

Judgment Seat: The approver whose name is on the record, whether or not the gate gave them enough to judge.

Episode 1:

The Last Human Who Read It

The Last Human Who Read It

Mark S. Carroll
·
Sep 9
Read full story

👋 Last week’s closing promised an episode called The Button That Used to Be a Person. The research changed my mind. The person didn’t disappear. They clicked. Episode One budgeted the reviewer’s attention. Episode Two asked what the sender owes before review begins. This week is about the moment in between: the approval itself, and what it can and can’t tell anyone who reads it later.

Episode 2:

The Work Is Finished. The Handoff Isn’t.

The Work Is Finished. The Handoff Isn’t.

Mark S. Carroll
·
Sep 16
Read full story

Research Binder: the receipts (citations + source notes) are compiled in a PDF at the bottom of this post.

Empathy Engine is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

⚡ Pressure

AI can expand what arrives for approval, and the click that clears it records presence, not judgment.

🎯 Payoff

This issue gives you the Gate Honesty Check: five questions that test whether an approval gate can support the judgment its record implies.

Pressure Moment

In mid-September, a software engineer posted a short after-action note. He had approved a pull request an AI agent wrote for a system he designed. Tests passed. The diff looked reasonable. He merged it without fully tracing one change. Two days later, a background job quietly started double-processing records. It’s one person’s public account, from software, and it says nothing about how often this happens. What it shows is narrower and more useful. He did check things. The merge record would look identical whether he traced that change or not.

The log recorded a human. It didn’t record the check.

I’ve inherited plenty of “approved” decisions in enterprise environments where the approval was real, but the judgment behind it wasn’t recoverable. We could prove who signed and what moved forward. Reconstructing what they had actually checked sometimes required emails, meeting notes, and conversations with people who were still around. The problem wasn’t necessarily the review. It was that the record couldn’t carry the review forward.

None of this means the approval was empty. Sometimes a green checkmark is the compact end of an excellent review. Sometimes the review was thinner. Sometimes we simply can’t tell, and not knowing is different from knowing that no judgment occurred. An incomplete record is not proof of an empty mind. The useful question is narrower: what evidence justifies treating this particular approval as meaningful oversight?

Start by separating the three questions a single click tends to collapse into one.

Was a human present? Did meaningful judgment occur? Did it improve the outcome? The log can answer the first: who clicked, when, and what state the system recorded. The other two need more evidence: what the person inspected, what evidence they used, and what would have made them refuse.

The one number on that graphic shows why the design around a decision matters. In a 199-participant experiment (Buçinca, Malaya & Gajos, 2021), people using simple explainable AI accepted incorrect advice on one wrong-advice subdecision about 64 percent of the time, versus 48 percent under cognitive-forcing designs. That’s a conditional result, not a 16-point cut in overall failure. It still shows that how a decision is set up changes what people do with bad advice. A human’s presence alone doesn’t guarantee that.


What the Record Can Prove

Organizations like signatures because one field compresses hours of work into “Approved by Taylor Kim.” That compression is useful. It also loses a lot.

A meta-analysis of 106 experiments and 370 effect sizes (Vaccaro, Almaatouq & Malone, 2024) found that human-AI combinations beat humans working alone on average, yet underperformed the better of the two working alone, with results varying widely by task. That finding supports neither “humans make AI better” nor “get humans out of the way.” Performance depends on the task, the system, and how the collaboration is designed. So read the record for what it can support: a sign-off shows presence, a review trail offers observability, and only outcome data speaks to effect.

The gap between a record and what actually happened isn’t new, and it isn’t an AI problem. There’s useful evidence from a setting with no AI at all.

In one hospital audit (Salgado, Barber & Danic, Journal of Patient Safety), independent audio review of 100 surgical cases found lower checklist fidelity than the paper completion records indicated. That doesn’t mean the clinicians were lying, or that no safety work occurred. The audit compared spoken checklist items with paper records, not anyone’s private judgment. Missing detail doesn’t prove missing judgment. It proves the record can’t settle the question on its own.

The answer isn’t more paperwork or surveillance. It’s deciding what each record is for. Approval status tells you who acted. Review evidence tells you what was inspected. Outcome data tells you what happened next. One checkmark can’t do all three jobs.


Watching Isn’t Stopping

Now imagine the reviewer spots a problem. What happens next? Darren Kimura, a technology executive quoted in CIO this month, described many companies that claim a human in the loop as actually having “a human watching the loop”: someone who can flag a concern but can’t stop, change, reject, or escalate it. He’s a vendor voice, so treat the line as a sharp distinction, not a measurement.

Share

The NIST AI Risk Management Framework tells organizations to assign mechanisms to supersede, disengage, or deactivate AI systems that perform inconsistently with their intended use. That’s a governance principle, not proof of outcome benefit. The closest outcome evidence comes from hospitals. A Cochrane review of escalation systems (McGaughey et al.) covered 4 randomized trials with 455,226 participants plus 7 nonrandomized studies. It couldn’t isolate authority as the active ingredient, because pathway design, recognition, communication, response, and implementation were bundled together.

That doesn’t mean authority is irrelevant. It means we should be precise about the pathway. Can the reviewer detect the problem? Can they act on it? Will the system honor the action? An approver without a usable pathway may be an observer with a better title.


Who Decided the Depth?

Episode One dealt with review capacity. This section asks a different question: when more work arrives, who decides how deep each review goes?

User's avatar

Continue reading this post for free, courtesy of Mark S. Carroll.

Or purchase a paid subscription.
© 2026 Mark S. Carroll · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture