Empathy Engine

Empathy Engine

The Work Got Better. What Did We Learn?

Who Trains the Next Reviewer? (JL-005 | The Judgment Layer)

Mark S. Carroll's avatar
Mark S. Carroll
Oct 07, 2026
∙ Paid

The Work Passed. Who’s Ready to Review?

Plain-English premise: AI can make today’s work better without telling you whether the person doing it is getting better. That second question needs its own evidence.

Judgment Seat: The future reviewer: the newer colleague whose judgment is still forming, and the lead deciding when to trust it.

👋 Last week’s closing promised an episode about the tasks that once formed future judgment, and what happens when they’re the first ones automated. The research changed my mind again. It didn’t establish that AI is quietly removing the practice that turns a newcomer into a reviewer. It showed something more useful: better work with AI and a more capable person without it are two different results, and the way the help works shapes which one you get.


Research Binder: the receipts (citations + source notes) are compiled in a PDF at the bottom of this post.

Empathy Engine is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

⚡ Pressure

AI can finish the first pass before a newer colleague has learned how to judge one.

🎯 Payoff

This issue gives you the Reviewer Bench Inventory: six questions that locate where newer colleagues practice judgment, get feedback, and get checked without the tool.

Pressure Moment

In late August 2026, Business Insider and Forbes reported that Valon, an AI startup, told most of its new hires to learn the job without AI until their manager was confident they could tell when the AI was wrong. Engineers were exempt, because all of their code already goes through peer review. According to the reporting, new hires had started asking tenured colleagues for help again.

That’s one company’s account, with no outcome data, and nothing in this issue supports restricting AI for newer people. What interests me is the test inside the policy. It didn’t ask whether new hires could produce good work with AI. It asked whether they could recognize when the AI was wrong, and it treated peer review as a place that question already gets answered.

I’ve stepped into a review and rewritten a backlog item myself because we needed to keep moving. The acceptance criteria got clearer, and the team had something it could use. What I hadn’t checked was whether the person who wrote it understood why I’d made those changes. I’d helped finish the work, but I’d left their learning unchecked.


Three Windows on the Same Work

When the work improves, it’s natural to assume the person doing it improved too. That used to be a safe assumption, because the draft improved as the person drafting it improved. AI can separate the two, so every result in this issue answers one of three questions.

With AI, how well did the work perform? Without AI, what could the person do independently? After a delay, what capability remained? A successful output answers the first question. It doesn’t answer the other two, and when a team collects only output evidence, its development claim remains unchecked. Count improved work. Check what people can do independently.

This is where the episode separates from Episode Two, which asked whether the sender can explain the work at handoff. That’s still the right gate, but an explanation at handoff is evidence about that work on that day. It isn’t evidence that judgment is forming, and an explanation can be generated too.

The louder version of this story runs on hiring data: Stanford’s Digital Economy Lab found US employment for 22-to-25-year-olds in the most AI-exposed occupations about 19% below comparable less-exposed work, mostly through slower hiring, while a Danish study found no such effect. Those numbers describe who gets onto the bench, not what anyone on it is learning.


The Output Gain Is Real

Start with the first window, where the news is good. You’ve seen “Generative AI at Work” (Brynjolfsson, Li & Raymond, QJE, 2025) in earlier episodes; it’s back for a different reason. In a staggered rollout of an AI assistant to 5,172 customer-support agents, 89% of them working outside the US, resolutions per hour rose 15.2% across all agents and 36% in the lowest skill quintile, a subgroup of the same sample.

Every bar measures performance with AI assistance, though the study does hint at learning. Agents with AI reached in about two months the output untreated agents needed eight to ten months to reach, and during rare outages they stayed faster than their pre-AI baseline. The authors describe those outage results as consistent with durable worker learning.

That’s a real signal, but not a controlled retention test: outages were rare and may not have been random. The output gain is real. Retained capability is a separate question, so count the productivity gain and assess development separately.


The Task Ended. Understanding Still Had to Be Tested.

The second window opens when the AI leaves the room. Shen and Tamkin, “How AI Impacts Skill Formation,” a 2026 preprint by two Anthropic researchers, recruited 52 programmers from a crowd-worker platform, all new to Trio, an unfamiliar Python library. Everyone got two coding tasks and up to 35 minutes. Half could use a GPT-4o chat assistant and half couldn’t, and then everyone took an immediate quiz with AI not permitted.

On average, the AI-assisted group scored 4.15 points lower on the 27-point quiz (d = 0.738, p = 0.010) and wasn’t significantly faster. Three cautions travel with that number. The sample was experienced, not junior: 29 of the 52 had seven or more years of coding, and only two per arm had one to three years. The quiz was immediate, so it can’t say what stuck. And it’s an Anthropic-affiliated preprint.

Its six usage patterns, where conceptual questioners outscored delegators, came from unrandomized groups of two to seven people, so they’re a clue, not a training method. Check understanding after the task, and keep the conclusion tied to the test.


Then the Evidence Pointed the Other Way

If the coding study were the whole story, this would be a warning and nothing more. Contractor and Reyes, “Experimental Evidence on the Learning Impact of Generative AI” (2026 preprint), randomized a US undergraduate sample to learn with AI allowed or forbidden, then tested everyone with AI and other resources barred. 211 students attended the first session, and 204 returned about a week later to answer new test items.

The AI-allowed group was 6.7 percentage points ahead immediately after learning and 5.1 points ahead about a week later. There’s a caveat on the first number: rule violations during the immediate test increased in the AI arm, and the authors estimated they could explain up to about one-third of the initial gain. The two figures are separate between-group estimates, not a measure of how much each student retained.

This is one study and one week, with undergraduates in a lab, on a different task and population from the coding study, so the two shouldn’t be averaged into a slogan. It does show that AI-assisted learning can carry over, which is why the third window is worth opening: check what remains, and when.


The Help Changed. So Did the Result.

A third experiment adds a distinction within one setting: the design of the assistance changed the outcome. Bastani and colleagues (”Generative AI without guardrails can harm learning,” PNAS, 2025) randomized nearly 1,000 high-school mathematics students in Turkey by classroom. Students practiced with GPT Base, an unguided chat assistant; with GPT Tutor, which used teacher-written hints and guardrails designed to guide rather than hand over answers; or without AI. Everyone then sat an exam without AI.

Share

During practice, GPT Base students scored 48% higher than control and GPT Tutor students 127% higher. On the exam, GPT Base students scored 17% lower than control. GPT Tutor students showed no statistically detectable difference from control, and no demonstrated exam advantage either; the authors describe the guardrails as largely mitigating the harm. The biggest practice gain didn’t prove a learning gain.

One detail should bother anyone who picks tools by satisfaction scores: students preferred the unguided tool. It’s high-school math with no delayed result, so the percentages don’t transfer to a product team. The question does: examine how the help works, and check what people can do afterward.


Practice Needs the Right Support

If development isn’t automatic, what builds it? The nostalgic answer is to protect the grunt work, because it was the apprenticeship. Some of it was; some was repetition nobody ever looked at. Even review doesn’t teach by default: in a 2013 Microsoft study, Bacchelli and Bird coded only 12 of 570 code-review comments as knowledge transfer, though learning can happen in ways a comment code wouldn’t capture.

User's avatar

Continue reading this post for free, courtesy of Mark S. Carroll.

Or purchase a paid subscription.
© 2026 Mark S. Carroll · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture