Editorial Note: The practitioners described in this article are composite characters created to illustrate the underlying feedback problem; they do not represent specific individuals or institutions. Where studies, legal cases, or operational examples are discussed, the article distinguishes documented findings from the broader argument being developed. The term “calibration gap” is used here as a conceptual framework for thinking about differences in human judgment that can emerge when AI users receive different levels of feedback about the outcomes of their decisions.
Does AI Make You Worse at Your Job? The Feedback Problem Nobody Measures
One radiologist read the scan wrong. No one will ever tell her.
Megan Walsh reads images for an independent group in Ohio, rotating among three sites each week. Her software is the same model, same version, same company as the system used at a university medical center five hundred miles away.
Last fall she signed a normal reading on a forty-two-year-old woman. The AI's prompt looked confident. So did hers. She sent the report and opened the next case.
The woman went elsewhere for follow-up. Those images and that pathology report will never flow back into Megan's system. There is no readout conference. There is no outcomes reconciliation. Nobody is going to tap her on the shoulder.
So it will not be found. Her error rate is not zero. It is invisible.
Elena Voss uses the same tool inside a teaching health system. Every read enters a structured review queue; the department holds a monthly quality conference; her calls get matched against final pathology or follow-up imaging. When the AI gives her a wrong prompt she usually knows within days. The feedback is free, institutional, and costless to her reputation.
Same capability. Same diligence. Same subscription tier. The only difference is how long it takes for someone to tell her she was wrong.
That difference compounds. Elena gets hundreds of "you were wrong last time" signals a year. Megan gets zero. A decade later their trust thresholds for the same piece of software will have drifted in opposite directions: one has learned when to doubt it, the other has learned to trust it. Both reason from real experience. Only one of those experiences contains corrections.
The part nobody measures: false negatives never come back
The argument so far—"some people get feedback, some don't"—is half right, and it is the shallow half. The deeper problem is structural: the feedback Megan can receive covers only one kind of error by construction.
She calls something suspicious, the biopsy is benign—she may know within days. Fast, relatively unambiguous, and usually tied to the same episode of care.
She calls something normal, and there is cancer—the patient may go elsewhere, get diagnosed much later, and the outcome may never return to her system.
Errors that trigger immediate follow-up are easier to feed back than errors that only become visible through later outcomes.
Her experience can therefore accumulate more lessons in the direction of "the AI is too conservative" and "that alert was a false alarm," while some of the lessons in the other direction never arrive. Her trust threshold can drift toward greater credence without any conscious decision to trust the system more. She is not necessarily fooling herself. She is learning inside an asymmetric feedback structure.
This is not a new insight dressed up in AI clothing. Diagnostic-accuracy research has a related problem with decades of literature: verification bias, also called work-up bias or partial verification. It arises when whether a patient's true condition gets established depends on the result of the test itself. In partial verification, positive tests are more likely to receive the reference standard while negative tests are less likely to be fully verified, which can bias estimates of diagnostic performance. Statistical methods exist to correct for this problem, but they depend on assumptions about the missing verification process.
The mechanism in this essay is not the same statistical bias. It is an operational analogue: the information needed to correct a person's judgment is more likely to return when an error is easy to observe than when the error remains hidden until a later outcome. Verification bias is primarily a problem of measuring diagnostic accuracy; the calibration gap proposed here is a problem of how experience changes the person using an AI system. That distinction is the shift that matters.
Before AI triage, a radiologist still looked at every screening examination, even though some missed cancers might only become visible through later follow-up. AI-supported workflows can change that allocation of attention. In Sweden's MASAI trial, for example, AI was used to triage mammograms: lower-risk examinations were assigned to single reading, while higher-risk examinations received double reading with AI support. The system therefore changed not only what the AI predicted, but also how much human review each case received.
That creates an important feedback question. If an error occurs in a case that receives less scrutiny, the signal may arrive only later through follow-up or an interval cancer. The system can still measure this at the population level, but the individual reader may never receive the case-specific correction.
This is why the MASAI results are relevant to the argument—but not because they prove that AI hides its own false negatives. The final analysis found that interval cancer rates were non-inferior with AI-supported screening compared with standard double reading, while the study also reported a substantial reduction in screen-reading workload. Those results demonstrate that system-level accuracy can be evaluated even when some errors are only discovered through later outcomes. They do not, by themselves, tell us whether an individual radiologist received the feedback needed to recalibrate her own judgment.
That distinction is the point. A health system can measure the accuracy of the workflow while the individual inside the workflow may still have incomplete information about her own errors.
The legal version runs on the same mechanism, skewed a different way. A fabricated citation gets caught by opposing counsel in a reply brief and teaches Ben Harris that the AI invents cases. It does not teach him which invented cases do not get caught. What he typically learns instead is overcorrection: stop using it for research, restrict it to formatting. He loses efficiency without gaining calibration, because he still has no idea what his escapes look like.
Which means the gap is worse than the initial description:
The problem is not that some people have feedback and others don't. It is that everyone's feedback loop is biased, and the bias always points toward trusting the AI more. The false-negative side only comes back if something outside the loop forces it home.
That is a testable prediction, not a piece of advice. It implies that any scheme relying on voluntary review will fail, because missed errors never volunteer to come home. It also explains why Rule 11 sanctions treat the symptom and not the disease: a sanction can only catch a fabricated citation that was already going to be discovered.
Two kinds of skill, only one grows on its own
To see why this is a calibration problem rather than a competence problem, separate two abilities that usually get lumped together.
The people who touch a model earliest and deepest develop an instinct for its edges. They know which questions are worth asking and which return nothing. This ability grows with exposure, the way an accent does. It is real, and it is not easily caught up with.
A second ability does not grow that way at all: knowing when you are probably wrong.
That one requires an external condition. You have to be in an environment where the truth arrives cheaply, where you make mistakes, and where something tells you about them.
Anyone who cooks already knows the difference. Two people follow the same smart-recipe app through the same steps. One tastes as she goes and learns within seconds whether it needs salt, at the cost of a minor adjustment. The other finds out when the dish comes back half-eaten and politely praised. By then there is nothing to fix. Same tool, similar hands. The only difference is whether there is a check that is cheap, independent, and available mid-process.
Developers feel it more sharply. Run your tests locally after every few lines and you find out in five minutes, at the cost of retyping twenty lines. Find out at code review or in the paging alarm and you pay with half a module and a postmortem. Same IDE, same Copilot.
Feedback lag has always done this work. It used to be hidden behind professional gatekeeping—you only got your work reviewed once you were inside certain institutions. Now the gate is gone and anyone can pull the same output a specialist gets. For the first time, lag is a mass variable instead of an occupational one.
And there is an amplifier baked into the model itself. The readability of generated text is not a reliable proxy for its correctness—a fully invented paragraph can be grammatically flawless, tonally restrained, and perfectly formatted. "Reads like truth" is no longer a safe shortcut to truth. In a study spanning multiple models and more than two thousand participants, making a user's view explicit increased the likelihood that models would produce agreeable responses even when those responses were wrong. The more effectively you can get an AI to engage with your framing, the more important it becomes to separate agreement from accuracy. Prompting skill is not a pure asset; its value still depends on whether the resulting output is independently checked.
So "use it more and you'll get better" is only half true. More use raises fluency and simultaneously increases the number of times you are persuaded by a fluent output. Without an independent verification source, the second quietly cancels the first.
The calibration gap: a definition
Call the divergence the calibration gap: systematic differentiation in judgment between people using the same tool at the same capability level, driven by differences in the cost and delay of truth feedback.
Only one variable changes: time to discovery multiplied by the cost of discovery.
The concept's sharpest feature is that it does not require cross-country comparison. Megan and Elena may work in different departments of the same hospital. Two litigators may work at opposite ends of the same city. Same model, same subscription, same release. The difference five years out comes from something very plain.
It also moves the discussion off the "personal AI literacy" target. Megan's disadvantage is not personal. Swap in any equally diligent person and the outcome does not change. The causes are specific and structural:
Data doesn't move. Medical imaging is scattered across thousands of systems that do not talk to each other. When a patient follows up elsewhere, the originating practice loses the outcome. This is not a technical problem. It is a property-and-incentives problem—no party has a reason to ship outcome data for free to support someone else's quality improvement.
Payment determines whether feedback is worth anything. In a per-procedure payment structure, the gains from quality improvement accrue to insurers and patients while the costs accrue to the physician and her clock. A quality conference consumes paid hours, and the avoided misdiagnosis cost does not hit her ledger. A large teaching hospital, by contrast, has a teaching budget, academic reputation, and residents to train; review is already priced into its cost structure.
The sample is too sparse for her to learn it herself. Screening errors are extremely sparse signals. A reader might see only a handful of false negatives in a year, scattered across years and separated by hundreds of correct calls. Human brains are bad at extracting patterns from sequences that are sparse, delayed, and noisy. Institutionalized retrospective review is doing for the reader something she cannot do biologically.
Why the market will not fix this
It is tempting to assume that if wiring the loop back were valuable, someone would sell it. That assumption fails for a clean reason: the data required to close the loop is fragmented across organizations whose incentives are not aligned.
The outcome that would tell Megan she was wrong lives in another organization's pathology system. Its value to her is enormous. Its value to the holder is zero or negative—it costs staff time to assemble, it creates liability exposure, and it improves a rival's quality metrics. No bilateral trade forms. Each pairwise negotiation is rational to decline, and the aggregate result is that the signal stays buried.
This is exactly the class of problem that can require architectures larger than an individual practitioner or a bilateral market: registries, shared data infrastructure, common standards, or institutions with enough authority and incentive to connect outcomes across organizations.
The broader lesson is not that higher AI adoption proves better feedback infrastructure. Adoption data cannot establish that causal relationship. The narrower point is more useful: the ability to close a feedback loop depends on institutional arrangements that can connect decisions to later outcomes, even when those outcomes occur somewhere else.
Penetration is easy to measure. The quality and completeness of the feedback loop behind that penetration are much harder to see.
The lawyer who found out six months later
Daniel Wu litigates at a big firm. He drafts and searches with AI, and before anything goes out someone runs the cases through Westlaw one by one. When it is wrong he knows before filing, at the cost of twenty minutes and some embarrassment.
Ben Harris practices alone on a flat-fee basis. His clients cannot tell the difference and nobody rechecks his briefs. He usually finds out he was wrong when opposing counsel writes "we could not locate these authorities" in a reply brief, or when a judge issues a sanction order. He finds out six months later, at the cost of a fine, an order to notify every court he misled, and a public record.
Lawyering exposes feedback lag in extreme form because the profession's feedback channels are intrinsically narrow. The great majority of federal civil cases never reach trial—they settle or they are dismissed. Settled cases produce no judicial ruling on which side's legal argument held up; clients pay, get a result, and generally cannot tell how many of the cases in their brief were real. A practitioner like Ben can go years without receiving a clear signal about the quality of his legal research.
Over the last few years, courts and tribunals have made one point increasingly clear in different contexts: AI use does not automatically displace existing responsibilities. Mata v. Avianca (S.D.N.Y., Judge Castel) involved attorneys who submitted fabricated authorities generated by ChatGPT and were sanctioned under Rule 11, including a $5,000 penalty and an order requiring notification of affected parties and judges. Moffatt v. Air Canada (2024 BCCRT 149) involved a different question: whether an airline could avoid responsibility for inaccurate information supplied by its chatbot by treating the chatbot as a separate legal entity.
The cases involve different actors and different legal questions, so they should not be treated as establishing one universal rule. But together they illustrate a narrower point relevant here: introducing an AI system does not, by itself, transfer responsibility for the consequences of its output away from the human or institution deploying or relying on it.
That sentence moves the question of who pays for the AI's mistake off the moral plane and onto the plane of responsibility allocation.
And do not assume the big-firm sign-off is solid. Under billing pressure partners glance. Juniors lean on seniors, seniors lean on process, process leans on templates. Institutional feedback is not truth; it is merely better in expectation. Closing the calibration gap is not accomplished by "instituting a review step." It depends on whether anyone in that step actually intends to find the error.
We already built things like this. Here is what they cost.
This is where the argument usually turns into a wish: "we need better institutions." But we do not need to imagine. Other fields have faced exactly the false-negative problem and built explicit machinery for it. The designs are instructive because they reveal both the solution shape and its failure modes.
Anesthesiology built the Closed Claims Project in the 1980s through the American Society of Anesthesiologists, with early work conducted in collaboration with the University of Washington. The idea was plain: malpractice claims provide a naturally occurring sample of serious failures, and systematic analysis can reveal injury patterns that routine monitoring may miss, especially rare ones. The project became an important source of evidence for patient-safety research, but its learning mechanism has a clear limitation: it depends on claims as the sampling frame. That means it is better at surfacing failures that generate claims than failures that never enter the claims system.
Notice the price tag on that learning: it required a national professional organization to maintain the program, years of accumulated cases, and a formal process for turning individual failures into shared safety knowledge. The architecture works precisely because the signal is collected beyond the individual practitioner.
Aviation went a different route with LOSA—Line Operations Safety Audit: trained peers observe routine operations, code threats and errors, and collect data under confidential, non-punitive conditions. It is not designed as a conventional compliance audit, and participation is voluntary. The design logic is explicit: if people expect observations to be used punitively, they may become less willing to participate or report what actually happened. The value of the feedback system depends on preserving enough trust for the observations to remain useful.
Read those two designs as a specification, because they are telling you something uncomfortable about what a real fix requires:
- Someone with authority has to compel the data, because no pairwise deal produces it (registry, mandatory linkage).
- The channel has to be insulated from blame, because the moment feedback carries reputational cost, people stop feeding it and start gaming it.
- It has to run on normal work, not on incidents, because failures are rare and you will wait years for a sample.
- It will still be slow, partial, and politically expensive.
Megan does not have any of these. She has a dashboard that tells her how many reads she completed today.
Objections that could break this framework
Put the strongest counterarguments on the table, ranked by lethality.
The only one that removes the foundation: engineering may actually drive the cost of correction close to zero. Tool use, self-checking, multi-agent cross-examination, automated evaluation—if that road works, calibration stops depending on the environment and becomes an engineering problem, and most of the argument above dissolves.
The honest state of play: there is neither proof that this road finishes nor proof that it doesn't. Three things are worth saying plainly. First, hallucination reappears at higher levels of reasoning—the errors you patch at the single-step level come back in new forms in multi-step inference. Second, evaluators fail too: using a model to grade itself produces systematic overestimation, and give it an objective and it will learn to game the grader rather than solve the problem. Third, and most troublesome: if the auto-evaluator and the evaluated model share a training distribution, they may make the same mistake, and you will never detect that collusion because you no longer have an independent reference. That last point is the mirror image of the essay's main claim. The same asymmetry that blinds a radiologist can blind an evaluation pipeline.
This is unresolved. Any claim offered with certainty here exceeds what the evidence supports.
Second: the absolute gain in fluency may far outweigh the loss in calibration. An experienced practitioner carrying a bit of overconfidence may still outproduce an extremely cautious novice. Observed effect sizes are generally small and mostly measured in lab settings.
Third: there is a version of this objection that is stronger than it looks—the better the AI gets, the more my framework should worry you, not less. If the tool is highly reliable, discovered errors become rarer, which means the feedback that sustains calibration arrives more slowly, which means trust drifts faster and in the direction of credence. Under this view the danger zone is not early, buggy AI. It is mature, dependable AI, right around the moment everyone stops doubting it. I do not have evidence for this beyond the logic; that is precisely why it belongs in the objection section. But it is the prediction I would test first.
Fourth: this is where the concept stops applying. Strategic decisions, creative writing, policy design—in these domains, when you get it wrong you often cannot tell when or by what standard. The calibration gap does not reach there. Talking about calibration in a domain with no ground truth is just dressing an intuition in technical clothing. That boundary needs to be marked.
So an honest statement of the position looks like this: if the cost of correction really gets engineered down to near zero, the calibration gap could become much less important and much of this argument would cease to apply. Until then, it remains a useful angle of observation.
You have to manufacture the signal
One consequence follows directly and it is the only part of this essay addressed to the individual reader. If the natural supply of corrective feedback is biased toward the errors that are easy to find, then waiting for nature to correct you is not a strategy—it is the mechanism of the failure. You cannot wait for the loop. You have to build a small one.
What that means in practice is not "be more careful." It is: introduce checks that do not depend on the world volunteering an error. Keep a fixed-size holdout of your own past outputs and re-verify them against primary sources at a fixed interval, regardless of outcome. Blind yourself to the AI's suggestion on a random sample of cases and compare. Seed your own workflow with known answers so you can see whether the tool—and you—are drifting. The common thread is that the verification must be independent, scheduled, and indifferent to whether anything went wrong. Spontaneous feedback teaches you about false positives. Only manufactured feedback teaches you about false negatives.
This is not a substitute for the institutional fix. It is a way of not being helpless while waiting for it. And it has a hard limit: no individual can manufacture the signal that lives in another organization's database. Megan cannot audit her own interval cancers. That is not a failure of diligence. It is a missing registry.
Back to Megan
She does not become more competent from reading this. What she needs is not a better model but a system that pushes that pathology report back onto her desk two years later. Until then she will keep signing reports, and she will never know which one she signed wrong.
All the cross-country comparison leads to a slightly unexpected place: the calibration gap is already splitting within the same city, the same company, the same office. National differences are second-order. Megan and Elena may be in different departments of the same hospital. Daniel and Ben may be at opposite ends of the same city. Same model, same subscription, same update. The difference five years out comes from something plain: how soon you find out you were wrong, and who bears the cost of finding out.
One question, worth more than any advice:
Do you have a verification source that is independent of the AI? If so, when was the last time it told you "you're wrong"—about something you would not otherwise have found?
That last clause is the whole essay. If the answer is "never," the problem is not the tool in front of you. The problem is that the loop has not been installed. And no product will install it for you.
Read More of Intelligenr
-
The Evolutionary Tree of AI: Why Transformers Dominated & What's Next
-
The Cognitive Trap Zone: Redesigning AI Tools Around Human Attention
References
-
Wan Nor Arifin & Umi Kalsom Yusof, 2022 Correcting for partial verification bias in diagnostic accuracy studies: A tutorial using R
-
Jessie Gommers et al., 2026 Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study
-
Lujain Ibrahim, Franziska Sofia Hafner & Luc Rocher, 2026 Training language models to be warm can reduce accuracy and increase sycophancy
-
Federal Aviation Administration, 2023 Line Operations Safety Assessments (LOSA)
Author Note: This article looks beyond whether an AI system is accurate to ask a different question: what happens to the person using it when the system's mistakes are difficult to see? As AI becomes embedded in professional decisions, the quality of the feedback surrounding the tool may matter as much as the tool itself.