The Expert Paradox in Radiology

Executive Research Report · Rapid Evidence Synthesis

The Expertise Paradox inRadiology

Experts make AI better, while novices may benefit most from it. A leadership operating framework for preserving diagnostic capability while adopting AI at scale.

+19.4 pp novice gain19.8% accuracy under wrong AI-6.0 pp unaided detection2.6× liability odds
The Expertise Paradox in Radiology — report cover
Executive Brief

Real gains, conditional on the human system that produces them

Artificial intelligence can expand differential diagnosis, reduce search burden, and support safer interpretation. Those gains are real. They are also conditional on the reader, the case, the interface, the sequence, model error, workload, and the learning context.

Novice diagnostic gain
+19.4 pp
Top-3 accuracy gain for neurology/neurosurgery residents with LLM assistance
Largest human benefit
Expert diagnostic gain
+4.4 pp
Neuroradiologists gained least, yet supplied the input that made the model best
Smallest benefit
Accuracy under wrong AI
19.8%
Correct BI-RADS by inexperienced readers when AI advice was incorrect (from 79%+)
Automation bias
Unaided detection shift
-6.0 pp
Non-AI adenoma detection after routine AI exposure in an adjacent specialty
Sentinel signal
The expertise paradox

Experts may contribute more to AI performance than they receive from it, while less experienced clinicians may receive larger immediate gains. If expert attention, teaching, and independent review are reduced because the model appears capable, the system may consume the very expertise that keeps the model useful.

What executive leadership should decide now

Six decisions that convert evidence into an operating posture

DecisionRecommended positionExecutive reason
1. Define protected cognitionRequire an independent first-pass interpretation before AI reveal for selected high-consequence, educationally sensitive, or difficult-to-recover tasks.Preserves retrieval effort, confidence calibration, and a defensible record of human reasoning.
2. Monitor human performanceTrack unaided performance, human-AI discordance, override quality, rare-case exposure, and AI-off recovery alongside model accuracy.A stable model does not prove a stable clinical team.
3. Segment delegationUse task consequence and recoverability to determine what may be automated, assisted, human-first, or restricted.Risk attaches to the work unit, not to AI as a single category.
4. Protect learning exposureGuarantee graded autonomy, case diversity, rare-event simulation, outcome feedback, and explicit teaching on AI failure modes.Efficiency gains can otherwise be purchased by consuming the future workforce’s learning opportunities.
5. Test resilienceRun scheduled AI-off drills and maintain downtime work standards for high-acuity imaging.Clinical capability must remain available during outages, drift, cyber disruptions, or novel presentations.
6. Assign an accountable ownerName a clinical executive responsible for the combined performance of model, workflow, users, and education.Distributed governance without decision rights produces invisible accountability gaps.
Board-level question

If the AI disappeared tomorrow, could this service line still deliver safe performance, recognize novelty, recover from error, and train the next generation? If the answer is unknown, the organization has an unmeasured resilience risk.

The executive thesis

Six propositions

AI assistance is neither uniformly beneficial nor uniformly harmful. Its effect depends on the reader, case, interface, sequence, model error, workload, and learning context.

Expertise is not a static credential. It is a perishable capability sustained by exposure, deliberate practice, feedback, metacognition, and error recovery.

The first cognitive event matters. Showing AI before independent interpretation changes search, anchoring, accountability, and the evidentiary record of human judgment.

Trainees require both AI literacy and protected opportunities to build unaided pattern recognition and diagnostic reasoning.

Model monitoring is necessary but insufficient. Leaders must monitor the coupled human-AI system and preserve AI-off performance.

The appropriate response is not blanket resistance. It is calibrated delegation supported by measurement, education, and clear stop rules.

Key definitions

The working vocabulary of expertise risk

Deskilling
A measurable decline in diagnostic, interpretive, procedural, or decision capability associated with reduced practice, altered exposure, or overreliance on automation.
Upskilling inhibition
Failure to develop a capability because automation arrives before foundational expertise has been acquired.
Automation bias
A tendency to favor automated advice, including errors of commission from following incorrect advice and errors of omission from failing to act when automation misses a finding.
Diagnostic resilience
The ability to maintain safe performance during model error, outage, drift, cyber disruption, unusual presentation, or other conditions outside routine AI support.
Expertise debt
A latent organizational obligation that arises when current productivity is achieved by reducing the learning and practice conditions required for future independent capability.
Section 1.1 · The experience gradient

Performing best when needed least

The 2026 brain MRI study by Schramm and colleagues provides the clearest recent radiology formulation of the paradox. Twelve readers evaluated 40 brain MRI examinations with confirmed diagnoses; large language models generated differentials from reader-authored findings, then readers revised their own diagnoses after reviewing GPT-4.1’s suggestions.

Top-3 accuracy before and after GPT-4.1

Interactive rebuild of Figure 2. Hover a bar for the absolute percentage-point gain.

Source: Schramm et al. (2026), Radiology, doi:10.1148/radiol.253477. Values are study-specific. pp = percentage points.

Figure 2 — mean top-three diagnostic accuracy before and after GPT-4.1 assistance
Figure 2. Mean top-three diagnostic accuracy before and after GPT-4.1 assistance in a 40-case brain MRI multireader study.

The gradient itself: benefit falls as experience rises

Diagnostic gain from assistance, plotted against reader experience.

Reader experience had a negative association with diagnostic benefit (beta = -0.10; P = .005) and a positive association with the correctness and completeness of the supplied findings.

Executive interpretation

Model accuracy was highest when neuroradiologists supplied the findings (78.8%–83.8% across models) and lower with resident input. Human benefit ran in the opposite direction. Experts contribute more to AI performance than they receive from it; the system depends on an expertise it can inadvertently deplete.

Experts make AI better

The quality of the human-supplied findings sets the ceiling on model performance. Remove expert input and the model degrades.

Novices gain most

Least-experienced readers saw the largest accuracy gains, which is why AI feels most valuable exactly where independent skill is least formed.

The trap

If the model looks capable, expert review and teaching are quietly withdrawn, consuming the input that made the model useful in the first place.

Section 1 · What the evidence says

Averages conceal the reader, case, and workflow that produce benefit or harm

AI improves group averages. It does not improve every reader, and when the model is wrong, experience provides partial protection, not immunity.

When the model is wrong, experience is not immunity

Correct BI-RADS assignment by reader experience, when AI advice was correct vs. incorrect (Figure 3).

Source: Dratsch et al. (2023), Radiology, doi:10.1148/radiol.222176. Prospective experiment, 27 radiologists.

Figure 3 — correct BI-RADS assignment when AI suggestions were correct versus incorrect
Figure 3. Correct BI-RADS assignment when AI suggestions were correct versus incorrect, stratified by reader experience.

Average benefit conceals individual harm

Heterogeneous treatment effects; conventional reader characteristics did not reliably predict who benefits (Figure 4).

Source: Yu et al. (2024), Nature Medicine, doi:10.1038/s41591-024-02850-w. Subgroup means; direction is illustrative.

Figure 4 — heterogeneity in AI treatment effects and the weak predictive value of reader characteristics
Figure 4. Heterogeneity in AI treatment effects and the weak predictive value of conventional reader characteristics.
Do not deploy by credentials alone

Seniority, subspecialty, and prior AI familiarity are not adequate proxies for safe collaboration. Local monitoring must examine individual and task-level patterns, while avoiding punitive surveillance or simplistic ranking.

Section 1.4 · Evidence matrix

The studies that most directly inform this report

StudyDesign & sampleFindingLeadership useKey limitation
Schramm et al., 2026Multireader; 12 readers; 40 brain MRI cases; 3 LLMsLLM benefit decreased with experience; LLM accuracy increased with expert-generated findings.Preserve expert input and protect novice learning.Single center, small sample, text-mediated workflow.
Dratsch et al., 2023Prospective; 27 radiologists; 50 mammogramsIncorrect AI advice degraded BI-RADS assignment across all experience groups.Train against plausible AI error; monitor agreement.Purported AI, experimental setting.
Yu et al., 2024140 radiologists; 324 cases; 15 chest X-ray tasksAI effects varied; experience did not reliably predict benefit.Avoid one-size-fits-all, credential-based deployment.Retrospective, controlled conditions.
Budzyn et al., 2025Multicenter observational; 1,443 non-AI colonoscopiesUnaided adenoma detection fell 6.0 pp after routine AI exposure.Measure post-adoption unaided performance.Adjacent specialty; residual confounding.
Keshavarz et al., 2025/26Systematic review; 14 training studies13 reported improved knowledge or performance; risks from overreliance and weak feedback.AI can upskill with diverse cases and mentorship.Heterogeneous interventions.
Bernstein et al., 2026Randomized mock-juror study; 282 participantsPlaintiff support lower when the radiologist interpreted independently first.Sequence and documentation shape perceived accountability.Hypothetical case; not the legal standard.
Section 1.5 · Two extremes the evidence does not support

Between alarmism and complacency

“AI will inevitably deskill radiologists.”

Direct longitudinal radiology evidence is limited, and AI-focused education can improve knowledge and performance. Deskilling is a measurable risk produced by specific task and workflow designs, not an unavoidable consequence.

“A human in the loop guarantees safety.”

Humans can anchor on incorrect AI, miss findings the model misses, or act as nominal reviewers. Oversight must be designed, trained, measured, and supported with time, authority, and feedback.

“Experts do not need AI safeguards.”

Incorrect AI affected very experienced readers, and credentials did not predict individual benefit. Safeguards should be task-sensitive and universal, with extra protection for trainees and high-consequence work.

“Higher combined accuracy proves durable value.”

Group averages can conceal individual harm, educational displacement, and fragility when AI is unavailable. Evaluate performance, distribution, independent capability, learning exposure, and resilience together.

Section 2 · How expertise can erode

Deskilling is not a single event; it is a set of mechanisms

Clinical expertise is a dynamic system output, sustained when clinicians retrieve knowledge, perform visual search, generate and rank hypotheses, confront uncertainty, receive feedback, and learn from error. AI can strengthen these processes or weaken them by presenting the answer too early, removing cases from review, homogenizing attention, or letting output volume displace reflection.

Section 2.1 · Seven mechanisms of expertise risk

What changes in the work, and the protective design

Retrieval displacement

AI supplies the finding or differential before the clinician constructs one.

Signal: shorter pre-AI dwell time; reduced independent differential breadth. Protect: independent first pass; delayed reveal; confidence recorded before AI.

Exposure displacement

AI triage or autonomous reads remove normal, subtle, or selected abnormal cases from human review.

Signal: changed case mix by role; declining rare-case contact. Protect: case-mix quotas; sampled review of AI-negative work; curated portfolio.

Search narrowing

Saliency cues and AI labels direct attention toward model-selected regions.

Signal: misses outside AI-marked areas. Protect: global search before reveal; interface sequencing; counterfactual cases.

Calibration decay

Clinician confidence becomes anchored to the model rather than to the quality of the evidence.

Signal: high confidence in concordant error; inappropriate override. Protect: pre-AI and post-AI confidence capture; feedback by agreement state.

Feedback attenuation

Speed and volume reduce peer review, outcome linkage, teaching, and reflection.

Signal: lower follow-up closure; fewer learning conferences per volume. Protect: automated outcome feeds; protected learning time; peer learning.

Recovery atrophy

As AI becomes always available, teams seldom practice without it.

Signal: slow or unsafe performance during outage; escalation failures. Protect: AI-off drills; downtime standards; drift simulation.

Role compression

The radiologist becomes a high-speed verifier of model output rather than an independent diagnostician.

Signal: lower consultation time; minimal edits; weak reasoning trace. Protect: task redesign that preserves synthesis and accountable judgment.

Section 2.2 · A sentinel longitudinal signal

Adjacent-specialty evidence that unaided performance can move

Non-AI adenoma detection, before vs. after AI exposure

Interactive rebuild of Figure 6. The 2025 ACCEPT analysis across four Polish centers.

Source: Budzyn et al. (2025), Lancet Gastroenterol Hepatol, doi:10.1016/S2468-1253(25)00133-5. Observational; not proof of causality or direct radiology evidence.

Figure 6 — non-AI adenoma detection before and after routine exposure to AI-assisted colonoscopy
Figure 6. Non-AI adenoma detection before and after routine exposure to AI-assisted colonoscopy.
Why leaders should care before radiology proof arrives

The governance question is not whether deskilling has been proven everywhere. It is whether the organization can detect a loss of unaided capability early enough to intervene. Waiting for rare harm or multi-year certainty can convert an observable operational risk into a workforce liability.

Expertise debt

Like technical debt, removing independent reads, normal-case exposure, or teaching improves immediate throughput. The cost appears later as weaker supervision, narrower pattern libraries, slower recovery, and fewer clinicians able to validate the next model. Unlike a software defect, expertise debt cannot be repaired instantly.

Proposed conceptual relation

Durable diagnostic expertise = independent exposure × deliberate reasoning × feedback quality × error recovery × calibrated AI use. The multiplicative form is conceptual, not a validated equation. It emphasizes that a near-zero value in any one domain can weaken the whole system.

Section 3 · Tomorrow’s workforce

AI literacy and diagnostic expertise are complementary, but must be sequenced

The educational evidence should prevent a one-sided narrative. Well-designed AI training can improve knowledge and performance; inadequate case diversity, feedback, and mentorship create countervailing risks.

AI-focused training can upskill trainees

Median performance before and after AI-focused radiology training (Figure 7).

Source: Keshavarz et al. (2025/2026), Academic Radiology, doi:10.1016/j.acra.2025.10.049. Summarizes 5 of 14 heterogeneous studies; not a pooled effect.

Figure 7 — selected median performance measures before and after AI-focused radiology training
Figure 7. Selected median performance measures before and after AI-focused radiology training.

Learning with AI vs. learning from AI

Learning with AI accelerates practice and provides immediate prompts. It produces short-term task support.

Learning from AI requires the trainee to understand why a suggestion is correct, when it fails, what evidence changes the diagnosis, and how to perform when the tool is absent. This builds transferable expertise.

A simulator that always announces a lesion can teach a search strategy that fails in ordinary clinical work.

Section 3.2 · Stage AI exposure by developmental need

From foundational trainee to clinical leader

Developmental stagePrimary learning needRecommended AI sequenceRequired evidence of competence
Foundational traineeVisual search, normal anatomy, pattern library, differential construction, uncertainty awareness.Substantial protected unaided interpretation; AI revealed after first pass; faculty review of discordance.Unaided performance, search completeness, confidence calibration, explanation of failure modes.
Intermediate traineeIntegrate findings with clinical context, manage ambiguity, learn prioritization.Mixed workflow by task; independent-first for high-consequence cases; guided AI comparison.Appropriate use, override reasoning, rare-case exposure, recovery during AI-off simulation.
Senior trainee / fellowSpeed without premature closure, consultation, supervision, system stewardship.Realistic AI-assisted workflow with periodic blinded and AI-off assessment; peer teaching.Stable combined and unaided performance, escalation judgment, teaching capability.
Practicing radiologistMaintain subspecialty expertise, adapt to tools, detect drift in self and model.Task-segmented AI use; longitudinal calibration feedback; skill-sensitive audit.Performance by agreement state, outcome closure, continuing education, downtime readiness.
Clinical leaderGovern task design, equity, learning systems, and organizational resilience.Access to model and human performance dashboards; authority to change sequence or pause use.Documented decisions, policy compliance, corrective action, evidence of capability preservation.
Avoid a false tradeoff

Delaying all AI until late training leaves residents unprepared for contemporary practice. Introducing unrestricted AI before foundational reasoning is established can inhibit skill formation. The answer is staged exposure with explicit independent-performance gates.

Section 4 · The Radiology Expertise Preservation Model

Five operating domains that make oversight a measurable system

Exposure provides the raw cases. Reasoning converts exposure into hypotheses. Feedback corrects and enriches them. Calibration aligns confidence and reliance with the evidence. Resilience keeps capability available when routine support fails. Governance supplies the owner, policy, data, time, and psychological safety across all five.

DomainDesign objectiveExample controlsFailure signal
ExposureMaintain variety, rarity, difficulty, normal comparators, and graded autonomy.Case-mix monitoring; sampled review of AI-negative cases; curated rare-case sets; equitable rotation.Stable volume with narrowing diversity or declining rare-event contact.
ReasoningPreserve independent search, synthesis, and uncertainty representation when consequences warrant.Delayed AI reveal; pre-AI findings and confidence; structured differential; protected time.Minimal independent record; immediate acceptance; high concordant error.
FeedbackConnect decisions to outcomes and make discordance educational.Pathology and follow-up linkage; peer learning; automated discrepancy feeds; coaching.Repeated error without learning closure; low follow-up completeness.
CalibrationMatch reliance to model, task, case, and reader conditions.Agreement-state metrics; override appropriateness; confidence shift; failure-mode simulations.Overreliance, underreliance, or wide individual variability without intervention.
ResilienceMaintain safe operation during outage, drift, novelty, and cyber disruption.AI-off drills; downtime worklists; manual prioritization; escalation rehearsal; recovery time objective.Performance collapse, backlog escalation, or unsafe workarounds when AI is absent.
Section 4.3 · Oversight mode should match risk

Which oversight design keeps the human engaged enough to catch failure?

Oversight modeHuman roleExpertise implicationAppropriate control
Human in the loopAction required for every case.Exposure may be preserved, but verification can still become passive.Independent-reasoning sample, agreement-state monitoring, adequate review time.
Human in parallelHuman and AI review independently, then reconcile.Strongest independent-performance signal; highest labor requirement.Use selectively for high-consequence validation, training, or drift investigation.
Human on the loopAI acts; humans supervise and intervene.Rare intervention can weaken recovery and vigilance.Alert-quality monitoring, simulation, minimum-intervention practice, escalation authority.
Human over the loopHumans govern the system, not individual outputs.Frontline exposure may disappear; governance expertise becomes critical.Robust surveillance, audit sampling, multidisciplinary owner, rapid rollback.
Human out of the loopAI performs and acts without contemporaneous review.Maximum exposure, displacement, and dependency risk.Restrict to evidence-supported, recoverable tasks with rigorous validation and contingency.
Section 4.1 · Delegate tasks, not professions

Risk attaches to the work unit, not to AI as a single category

A single examination contains tasks in different risk categories. Automated segmentation may be highly recoverable; autonomous exclusion of a subtle high-consequence diagnosis may not be. Use consequence and recoverability to place each task.

Task delegation calculator

Set the clinical consequence and recoverability of a specific task to see the recommended delegation mode. This mirrors the executive delegation matrix; example placements are illustrative and require local clinical, regulatory, legal, and technical review.

How much harm results if this task is done wrong?

MinimalModerateSeriousCritical

How easily can an error be detected and reversed before patient harm?

Very hardHardModerateEasy
Recommended mode
Human-First
Higher consequence · recoverable
Human-First
Section 4.1 · Task policy elements

Seven decisions to specify per task

Task policy elementRequired decision
Intended useWhat exact decision or work product will AI influence, for which population, modality, site, and user?
SequenceWill AI be visible before, during, or after independent interpretation?
Human authorityWho may accept, override, escalate, or suspend the output, and what evidence is required?
RecoveryHow can an error be detected and reversed before patient harm?
Learning sensitivityDoes automation remove repetitions, case exposure, or graded autonomy needed for competence?
MonitoringWhich model, human, combined, equity, educational, and resilience measures will be reviewed?
Stop ruleWhat threshold triggers investigation, sequence change, retraining, rollback, or suspension?
Design principle

The safest workflow is not always the one with the most human steps. It is the one that preserves the right human capability at the point where error is consequential and difficult to recover.

Section 5 · The measurement system

Monitor the model, the human, the team, the learning environment, and AI-off capability

Model monitoring is indispensable but insufficient. It cannot reveal whether clinicians are becoming dependent, whether trainees are losing exposure, or whether combined performance varies across readers. Treat human-AI collaboration as a coupled clinical process.

Measure four performance states

Typical monitoring focuses on AI-alone and combined performance; expertise preservation demands all four.

Author-developed comparison of monitoring focus. Illustrative coverage, not measured values.

The four states

Human alone

What capability exists without the tool? Blinded sample, simulation, delayed reveal, or AI-off drill.

AI alone

What can the model do on local data now? Site validation, silent mode, drift and subgroup monitoring.

Human plus AI

Does the combined workflow improve patient-relevant performance? Prospective comparison, agreement-state review.

Human after interruption

Can the service recover during outage, novelty, or drift? Scheduled downtime simulation and after-action review.

A stable model can coexist with a changing human system

If model sensitivity remains constant while unaided clinician sensitivity declines, the dashboard may stay green until the model is incorrect or unavailable. Monitoring only the algorithm can create a false sense of assurance.

Section 5.3 · Agreement states

The minimum analytic unit is the four-state agreement grid

Human correct / AI correct

Is speed gained without weakening independent search? Sample for premature closure and opportunities to simplify low-risk workflow.

Human correct / AI wrong

Did the human resist the error, and what evidence supported the override? Reinforce appropriate override; investigate model failure.

Human wrong / AI correct

Was useful AI information seen, understood, and acted on? Improve presentation, training, escalation, or confidence communication.

Human wrong / AI wrong

Do humans and models share a blind spot? Escalate to failure-mode analysis, education, workflow redesign, or additional safeguards.

Section 5.2 · Expertise-preservation scorecard

What to review, and how often

DomainMeasureReview cadenceGuardrail question
QualityCombined diagnostic performance (sensitivity, specificity, critical miss rate).MonthlyIs combined performance improving for all material subgroups?
Independent capabilityUnaided performance on blinded, delayed-reveal, or AI-off cases.QuarterlyIs human-alone capability stable within tolerance?
RelianceOverride appropriateness after adjudication.Monthly → quarterlyAre users over-relying, under-relying, or varying unpredictably?
CalibrationConfidence shift by agreement state.QuarterlyDoes AI increase confidence even when wrong?
ExposureLearning exposure index by rarity, difficulty, and autonomy.Monthly (trainees)Which cases have disappeared from human practice?
FeedbackOutcome closure rate.MonthlyAre decisions converted into learning?
ResilienceAI-off recovery performance.Semiannual drillCan the service sustain a safe degraded mode?
WorkforceCognitive burden and trust calibration.Baseline, 30/60/90 daysIs safety achieved through unsustainable hidden work?
EquityBenefit and harm distribution by site, shift, and role.QuarterlyWho gains, who bears verification work, who loses learning?
Measure for learning, not surveillance

Clinician-level data are necessary because AI effects vary by user, and they are sensitive. If dashboards rank productivity or punish disagreement, users conceal uncertainty and game the workflow. Default use should be system improvement and confidential development, with escalation reserved for validated, repeated safety concerns.

Section 6 · Governance and accountability

Expertise risk crosses many domains, so ownership must be explicit

A committee can coordinate evaluation, cybersecurity, privacy, procurement, equity, and monitoring, but it does not substitute for an accountable executive. Each use case needs a named clinical owner with authority to change sequence, require education, commission an audit, modify a threshold, suspend use, or initiate rollback.

Workflow sequence can carry legal meaning

Mock-juror responses in a hypothetical missed-bleed case, by radiologist-AI sequence (Figure 11).

Source: Bernstein et al. (2026), Nature Health, doi:10.1038/s44360-026-00085-2. Randomized study of 282 mock jurors. Behavioral research, not legal advice.

Figure 11 — mock-juror responses in a hypothetical missed-bleed case by radiologist-AI workflow sequence
Figure 11. Mock-juror responses in a hypothetical missed-bleed case by radiologist-AI workflow sequence.
The finding

When AI correctly flagged the CT, 74.7% of mock jurors sided with the plaintiff if the radiologist read once after seeing AI, versus 52.9% when the radiologist interpreted independently first. The odds of siding with the plaintiff were 2.6 times higher in the single-read condition.

Legal caution

Independent-first documentation is not a liability shield and may be inappropriate for some tasks. Obtain jurisdiction-specific legal review. The broader point: workflow architecture, training, alerts, and documentation can become part of the factual record when harm occurs.

Section 6.1 · Accountability map

Nondelegable accountability by role

RoleNondelegable accountability
Board / quality committeeConfirm that AI risk oversight includes human capability, resilience, and equity, not only vendor and cybersecurity risk.
Chief medical / clinical officerSet enterprise clinical AI principles, resolve cross-service risk, and ensure accountability for patient outcomes.
Radiology chair / executiveOwn task policy, staffing and education implications, performance review, and clinical escalation.
Program director / education leaderProtect graded autonomy, exposure portfolio, AI literacy, and independent competency assessment.
AI governance committeeStandardize intake, evidence review, validation, monitoring, change control, and incident response.
Quality & safety leaderIntegrate AI agreement states and expertise signals into peer learning, root cause analysis, and safety reporting.
Vendor / technology partnerProvide performance evidence, intended-use limits, version transparency, logging, update notice, and audit/rollback support.
Frontline userUse within policy, document material disagreement, report failure, participate in training, and retain authority to escalate.
Section 6.5 · Maturity model

From tool adoption to an expertise-resilient system

Level 1

Tool adoption. Model integrated; success defined by accuracy or speed. Expertise risk is assumed to be controlled by having a human reviewer.

Level 2

Clinical validation. Local performance and intended use evaluated; users trained. Automation bias recognized, but unaided capability and exposure are not measured.

Level 3

Coupled-system governance. Human, model, and combined performance monitored by agreement state and subgroup. Task policies and learning safeguards are active.

Level 4

Expertise-resilient system. Longitudinal capability, education, AI-off recovery, equity, and change control integrated. Expertise treated as a strategic safety asset.

Section 7 · A 90-day expertise-preservation pilot

Test one consequential use case, learn quickly, create a reusable standard

The first pilot should be clinically meaningful but bounded: an existing or near-term deployment, measurable outputs, accessible outcome data, a defined user group, and a plausible expertise-sensitive mechanism.

Section 7.2 · Pilot protocol

Six phases, each with analytic outputs

DAYS0-10
Charter

Name the owner; define the frame

Intended use, sequence, users, exclusions, escalation, stop rules, and data governance.

Outputs: signed use-case charter; risk register; metric dictionary; training plan.

DAYS11-30
Baseline

Measure the starting point

Current performance, case mix, confidence where feasible, feedback closure, workload, and downtime process.

Outputs: baseline dashboard; data-quality assessment; power and sampling plan.

DAYS31-45
Activation

Train and launch a protected trial

Train on model strengths, failure modes, overrides, reporting, and an independent-first workflow.

Outputs: competency evidence; technical validation; early burden review.

DAYS46-60
Stabilize

Review discordance weekly

Correct integration and interface failures; monitor safety and user burden.

Outputs: agreement-state analysis; incident log; process fidelity; corrective actions.

DAYS61-75
Resilience

Run rare-case and AI-off exercises

Assess recovery, escalation, and independent performance under simulated absence.

Outputs: simulation results; downtime performance; learning gaps; remediation plan.

DAYS76-90
Decision

Compare to baseline and guardrails

Evaluate subgroup, educational, legal, and operational implications.

Outputs: scale / revise / pause / stop recommendation; reusable policy and lessons learned.

Section 6.5 + 7.4 · Maturity self-check

How expertise-resilient is your current deployment?

Check each control your service line has in place. The meter maps to the report’s four-level maturity model.

0 / 8 controlsLevel 1
Level 1 · Tool adoption. Expertise risk is assumed to be controlled simply by having a human reviewer. Unaided capability and case exposure are not yet measured.

Named accountable clinical owner

One executive has authority to change sequence, pause use, or initiate rollback.

Task-level delegation policy

Each task is placed by consequence and recoverability, not labeled broadly as “AI-assisted.”

Independent-first workflow where warranted

Protected cognition captured before AI reveal for high-consequence, educationally sensitive tasks.

Unaided performance monitored

Human-alone capability tracked via blinded, delayed-reveal, or AI-off samples.

Agreement-state review

Cases classified into the four agreement states; overrides adjudicated for appropriateness.

Learning exposure protected

Case-mix quotas, rare-event simulation, and graded autonomy for trainees.

AI-off resilience drills

Scheduled downtime exercises with measured recovery, escalation, and backlog performance.

Board-visible guardrails and stop rules

Prespecified thresholds trigger investigation, sequence change, retraining, or rollback.

Do not force a single composite score

A workflow that improves combined sensitivity but materially worsens unaided performance, trainee exposure, or cognitive burden presents a leadership tradeoff, not an automatic success. Keep outcomes visible as a balanced set and document the decision rationale.

References

Selected literature and guidance

References are current through July 17, 2026. Filter by evidence stream; each links to the DOI, publisher, or authoritative source. This report is a rapid executive synthesis, not a registered systematic review.

Schramm, S., Le Guellec, B., Topka, M., et al. (2026).Reader study

Performing best when needed least: Reader experience shapes accuracy gains in LLM-assisted brain MRI differential diagnosis. Radiology, 319(2).

View DOI / source →

Dratsch, T., Chen, X., Rezazade Mehrizi, M., et al. (2023).Reader study

Automation bias in mammography: The impact of AI BI-RADS suggestions on reader performance. Radiology, 307(4), e222176.

View DOI / source →

Yu, F., Moehring, A., Banerjee, O., et al. (2024).Reader study

Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30, 837-849.

View DOI / source →

Budzyn, K., Romanczyk, M., Kitala, D., et al. (2025).Longitudinal

Endoscopist deskilling risk after exposure to AI in colonoscopy: A multicentre, observational study. Lancet Gastroenterol Hepatol, 10(10), 896-903.

View DOI / source →

Heudel, P. E., Crochet, H., Filori, Q., et al. (2026).Longitudinal

AI in medicine: A scoping review of the risk of deskilling and loss of expertise among physicians. ESMO Real World Data & Digital Oncology, 12, 100693.

View DOI / source →

Macnamara, B. N., Berber, I., Cavusoglu, M. C., et al. (2024).Longitudinal

Does using AI assistance accelerate skill decay and hinder skill development without the performers’ awareness? Cognitive Research, 9, 46.

View DOI / source →

Natali, C., Marconi, L., Dias Duran, L. D., et al. (2025).Longitudinal

AI-induced deskilling in medicine: A mixed-method review and research agenda. Artificial Intelligence Review, 58, 356.

View DOI / source →

Keshavarz, P., Mohammadigoldar, Z., Bedayat, A., et al. (2025/2026).Education

AI education in radiology training: A systematic review of effectiveness, barriers, and future directions. Academic Radiology, 33(3), 695-706.

View DOI / source →

Kitamura, F., Kline, T., Warren, D., et al. (2025).Education

Teaching AI for radiology applications: A multisociety-recommended syllabus from the AAPM, ACR, RSNA, and SIIM. Radiology: AI, 7(6), e250137.

View DOI / source →

Tejani, A. S., Elhalawani, H., Moy, L., Kohli, M., & Kahn, C. E. (2023).Education

Artificial intelligence and radiology education. Radiology: Artificial Intelligence, 5(1), e220084.

View DOI / source →

Tejani, A. S., Rauschecker, A. M., Kohli, M., Mongan, J., & Moy, L. (2026).Education

AI for radiology: A primer, part II. Interacting with AI results. Radiology, 320(1).

View DOI / source →

Brady, A. P., Allen, B., Chong, J., et al. (2024).Governance

Developing, purchasing, implementing, and monitoring AI tools in radiology: A multisociety statement (ACR, CAR, ESR, RANZCR, RSNA). Radiology: AI, 6(1), e230513.

View DOI / source →

Venugopal, V. K., Khubchandani, S. A., Liew, C. J. Y., et al. (2026).Governance

Postdeployment monitoring and surveillance methods, guidelines, and possibilities for AI in radiology. RadioGraphics, 46(7), e250173.

View DOI / source →

Doo, F. X., Tripathi, S., Huisman, M., et al. (2026).Governance

Alignment of policy, practice, and patient safety for trustworthy AI in radiology. Radiology: Artificial Intelligence.

View DOI / source →

Lekadir, K., Frangi, A. F., Porras, A. R., et al. (2025).Governance

FUTURE-AI: International consensus guideline for trustworthy and deployable AI in healthcare. BMJ, 388, e081554.

View DOI / source →

Joint Commission & Coalition for Health AI. (2025).Governance

Guidance on the responsible use of artificial intelligence in healthcare.

View DOI / source →

Bernstein, M. H., Sheppard, B., Bruno, M. A., et al. (2026).Legal & behavioral

The radiologist-AI workflow and the risk of medical malpractice claims. Nature Health, 1, 386-389.

View DOI / source →

Mello, M. M., & Guha, N. (2024).Legal & behavioral

Understanding liability risk from using healthcare AI tools. New England Journal of Medicine, 390(3), 271-278.

View DOI / source →

Parasuraman, R., & Manzey, D. H. (2010).Human factors

Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381-410.

View DOI / source →

Goddard, K., Roudsari, A., & Wyatt, J. C. (2012).Human factors

Automation bias: A systematic review of frequency, effect mediators, and mitigators. JAMIA, 19(1), 121-127.

View DOI / source →

Ghassemi, M., Oakden-Rayner, L., & Beam, A. L. (2021).Human factors

The false hope of current approaches to explainable AI in health care. Lancet Digital Health, 3(11), e745-e750.

View DOI / source →

McMillan, A. B. (2026).Perspective

The potential expertise paradox in AI-assisted radiology. Radiology, 319(2).

View DOI / source →

The leadership standard

Radiology spent the past decade asking whether AI can match human performance. The next leadership question is harder: what kind of human capability will the AI-enabled system produce over time?

The answer is not nostalgia for unaided practice. It is disciplined augmentation: protect independent cognition where error is consequential, preserve exposure where expertise is still developing, convert disagreement into feedback, measure human-alone and AI-off performance, and give a clinical leader the authority to intervene. AI should make radiology more capable, not merely more dependent.

The Expertise Paradox in Radiology — Executive Leadership Dashboard. Prepared by Kelly Emrick, DHSc, PhD, MBA, BSRT(ARRT)R. Interactive rebuilds of published figures are provided alongside the original figures; empirical claims are cited to their sources. Author-developed frameworks are conceptual and proposed for local validation.

Homekellyemrick.com