The Expertise Paradox inRadiology
Experts make AI better, while novices may benefit most from it. A leadership operating framework for preserving diagnostic capability while adopting AI at scale.

Real gains, conditional on the human system that produces them
Artificial intelligence can expand differential diagnosis, reduce search burden, and support safer interpretation. Those gains are real. They are also conditional on the reader, the case, the interface, the sequence, model error, workload, and the learning context.
Experts may contribute more to AI performance than they receive from it, while less experienced clinicians may receive larger immediate gains. If expert attention, teaching, and independent review are reduced because the model appears capable, the system may consume the very expertise that keeps the model useful.
Figure 1 · The Radiology AI Expertise Paradox
Visible near-term gains can coexist with latent long-horizon capability risk. The conceptual model is a leadership hypothesis for local testing, informed by Schramm et al. (2026), McMillan (2026), Dratsch et al. (2023), Yu et al. (2024), and Heudel et al. (2026).
Six decisions that convert evidence into an operating posture
| Decision | Recommended position | Executive reason |
|---|---|---|
| 1. Define protected cognition | Require an independent first-pass interpretation before AI reveal for selected high-consequence, educationally sensitive, or difficult-to-recover tasks. | Preserves retrieval effort, confidence calibration, and a defensible record of human reasoning. |
| 2. Monitor human performance | Track unaided performance, human-AI discordance, override quality, rare-case exposure, and AI-off recovery alongside model accuracy. | A stable model does not prove a stable clinical team. |
| 3. Segment delegation | Use task consequence and recoverability to determine what may be automated, assisted, human-first, or restricted. | Risk attaches to the work unit, not to AI as a single category. |
| 4. Protect learning exposure | Guarantee graded autonomy, case diversity, rare-event simulation, outcome feedback, and explicit teaching on AI failure modes. | Efficiency gains can otherwise be purchased by consuming the future workforce’s learning opportunities. |
| 5. Test resilience | Run scheduled AI-off drills and maintain downtime work standards for high-acuity imaging. | Clinical capability must remain available during outages, drift, cyber disruptions, or novel presentations. |
| 6. Assign an accountable owner | Name a clinical executive responsible for the combined performance of model, workflow, users, and education. | Distributed governance without decision rights produces invisible accountability gaps. |
If the AI disappeared tomorrow, could this service line still deliver safe performance, recognize novelty, recover from error, and train the next generation? If the answer is unknown, the organization has an unmeasured resilience risk.
Six propositions
AI assistance is neither uniformly beneficial nor uniformly harmful. Its effect depends on the reader, case, interface, sequence, model error, workload, and learning context.
Expertise is not a static credential. It is a perishable capability sustained by exposure, deliberate practice, feedback, metacognition, and error recovery.
The first cognitive event matters. Showing AI before independent interpretation changes search, anchoring, accountability, and the evidentiary record of human judgment.
Trainees require both AI literacy and protected opportunities to build unaided pattern recognition and diagnostic reasoning.
Model monitoring is necessary but insufficient. Leaders must monitor the coupled human-AI system and preserve AI-off performance.
The appropriate response is not blanket resistance. It is calibrated delegation supported by measurement, education, and clear stop rules.
The working vocabulary of expertise risk
- Deskilling
- A measurable decline in diagnostic, interpretive, procedural, or decision capability associated with reduced practice, altered exposure, or overreliance on automation.
- Upskilling inhibition
- Failure to develop a capability because automation arrives before foundational expertise has been acquired.
- Automation bias
- A tendency to favor automated advice, including errors of commission from following incorrect advice and errors of omission from failing to act when automation misses a finding.
- Diagnostic resilience
- The ability to maintain safe performance during model error, outage, drift, cyber disruption, unusual presentation, or other conditions outside routine AI support.
- Expertise debt
- A latent organizational obligation that arises when current productivity is achieved by reducing the learning and practice conditions required for future independent capability.
Performing best when needed least
The 2026 brain MRI study by Schramm and colleagues provides the clearest recent radiology formulation of the paradox. Twelve readers evaluated 40 brain MRI examinations with confirmed diagnoses; large language models generated differentials from reader-authored findings, then readers revised their own diagnoses after reviewing GPT-4.1’s suggestions.
Top-3 accuracy before and after GPT-4.1
Interactive rebuild of Figure 2. Hover a bar for the absolute percentage-point gain.
Source: Schramm et al. (2026), Radiology, doi:10.1148/radiol.253477. Values are study-specific. pp = percentage points.

The gradient itself: benefit falls as experience rises
Diagnostic gain from assistance, plotted against reader experience.
Reader experience had a negative association with diagnostic benefit (beta = -0.10; P = .005) and a positive association with the correctness and completeness of the supplied findings.
Model accuracy was highest when neuroradiologists supplied the findings (78.8%–83.8% across models) and lower with resident input. Human benefit ran in the opposite direction. Experts contribute more to AI performance than they receive from it; the system depends on an expertise it can inadvertently deplete.
Experts make AI better
The quality of the human-supplied findings sets the ceiling on model performance. Remove expert input and the model degrades.
Novices gain most
Least-experienced readers saw the largest accuracy gains, which is why AI feels most valuable exactly where independent skill is least formed.
The trap
If the model looks capable, expert review and teaching are quietly withdrawn, consuming the input that made the model useful in the first place.
Averages conceal the reader, case, and workflow that produce benefit or harm
AI improves group averages. It does not improve every reader, and when the model is wrong, experience provides partial protection, not immunity.
When the model is wrong, experience is not immunity
Correct BI-RADS assignment by reader experience, when AI advice was correct vs. incorrect (Figure 3).
Source: Dratsch et al. (2023), Radiology, doi:10.1148/radiol.222176. Prospective experiment, 27 radiologists.

Average benefit conceals individual harm
Heterogeneous treatment effects; conventional reader characteristics did not reliably predict who benefits (Figure 4).
Source: Yu et al. (2024), Nature Medicine, doi:10.1038/s41591-024-02850-w. Subgroup means; direction is illustrative.

Seniority, subspecialty, and prior AI familiarity are not adequate proxies for safe collaboration. Local monitoring must examine individual and task-level patterns, while avoiding punitive surveillance or simplistic ranking.
The studies that most directly inform this report
| Study | Design & sample | Finding | Leadership use | Key limitation |
|---|---|---|---|---|
| Schramm et al., 2026 | Multireader; 12 readers; 40 brain MRI cases; 3 LLMs | LLM benefit decreased with experience; LLM accuracy increased with expert-generated findings. | Preserve expert input and protect novice learning. | Single center, small sample, text-mediated workflow. |
| Dratsch et al., 2023 | Prospective; 27 radiologists; 50 mammograms | Incorrect AI advice degraded BI-RADS assignment across all experience groups. | Train against plausible AI error; monitor agreement. | Purported AI, experimental setting. |
| Yu et al., 2024 | 140 radiologists; 324 cases; 15 chest X-ray tasks | AI effects varied; experience did not reliably predict benefit. | Avoid one-size-fits-all, credential-based deployment. | Retrospective, controlled conditions. |
| Budzyn et al., 2025 | Multicenter observational; 1,443 non-AI colonoscopies | Unaided adenoma detection fell 6.0 pp after routine AI exposure. | Measure post-adoption unaided performance. | Adjacent specialty; residual confounding. |
| Keshavarz et al., 2025/26 | Systematic review; 14 training studies | 13 reported improved knowledge or performance; risks from overreliance and weak feedback. | AI can upskill with diverse cases and mentorship. | Heterogeneous interventions. |
| Bernstein et al., 2026 | Randomized mock-juror study; 282 participants | Plaintiff support lower when the radiologist interpreted independently first. | Sequence and documentation shape perceived accountability. | Hypothetical case; not the legal standard. |
Between alarmism and complacency
“AI will inevitably deskill radiologists.”
Direct longitudinal radiology evidence is limited, and AI-focused education can improve knowledge and performance. Deskilling is a measurable risk produced by specific task and workflow designs, not an unavoidable consequence.
“A human in the loop guarantees safety.”
Humans can anchor on incorrect AI, miss findings the model misses, or act as nominal reviewers. Oversight must be designed, trained, measured, and supported with time, authority, and feedback.
“Experts do not need AI safeguards.”
Incorrect AI affected very experienced readers, and credentials did not predict individual benefit. Safeguards should be task-sensitive and universal, with extra protection for trainees and high-consequence work.
“Higher combined accuracy proves durable value.”
Group averages can conceal individual harm, educational displacement, and fragility when AI is unavailable. Evaluate performance, distribution, independent capability, learning exposure, and resilience together.
Deskilling is not a single event; it is a set of mechanisms
Clinical expertise is a dynamic system output, sustained when clinicians retrieve knowledge, perform visual search, generate and rank hypotheses, confront uncertainty, receive feedback, and learn from error. AI can strengthen these processes or weaken them by presenting the answer too early, removing cases from review, homogenizing attention, or letting output volume displace reflection.
Figure 5 · The reinforcing expertise-erosion loop

A reinforcing loop with three points of interruption. Author-developed leadership hypothesis informed by Parasuraman & Manzey (2010), Goddard et al. (2012), Macnamara et al. (2024), Heudel et al. (2026), and McMillan (2026). Local testing is required.
What changes in the work, and the protective design
Retrieval displacement
AI supplies the finding or differential before the clinician constructs one.
Signal: shorter pre-AI dwell time; reduced independent differential breadth. Protect: independent first pass; delayed reveal; confidence recorded before AI.
Exposure displacement
AI triage or autonomous reads remove normal, subtle, or selected abnormal cases from human review.
Signal: changed case mix by role; declining rare-case contact. Protect: case-mix quotas; sampled review of AI-negative work; curated portfolio.
Search narrowing
Saliency cues and AI labels direct attention toward model-selected regions.
Signal: misses outside AI-marked areas. Protect: global search before reveal; interface sequencing; counterfactual cases.
Calibration decay
Clinician confidence becomes anchored to the model rather than to the quality of the evidence.
Signal: high confidence in concordant error; inappropriate override. Protect: pre-AI and post-AI confidence capture; feedback by agreement state.
Feedback attenuation
Speed and volume reduce peer review, outcome linkage, teaching, and reflection.
Signal: lower follow-up closure; fewer learning conferences per volume. Protect: automated outcome feeds; protected learning time; peer learning.
Recovery atrophy
As AI becomes always available, teams seldom practice without it.
Signal: slow or unsafe performance during outage; escalation failures. Protect: AI-off drills; downtime standards; drift simulation.
Role compression
The radiologist becomes a high-speed verifier of model output rather than an independent diagnostician.
Signal: lower consultation time; minimal edits; weak reasoning trace. Protect: task redesign that preserves synthesis and accountable judgment.
Adjacent-specialty evidence that unaided performance can move
Non-AI adenoma detection, before vs. after AI exposure
Interactive rebuild of Figure 6. The 2025 ACCEPT analysis across four Polish centers.
Source: Budzyn et al. (2025), Lancet Gastroenterol Hepatol, doi:10.1016/S2468-1253(25)00133-5. Observational; not proof of causality or direct radiology evidence.

The governance question is not whether deskilling has been proven everywhere. It is whether the organization can detect a loss of unaided capability early enough to intervene. Waiting for rare harm or multi-year certainty can convert an observable operational risk into a workforce liability.
Like technical debt, removing independent reads, normal-case exposure, or teaching improves immediate throughput. The cost appears later as weaker supervision, narrower pattern libraries, slower recovery, and fewer clinicians able to validate the next model. Unlike a software defect, expertise debt cannot be repaired instantly.
Durable diagnostic expertise = independent exposure × deliberate reasoning × feedback quality × error recovery × calibrated AI use. The multiplicative form is conceptual, not a validated equation. It emphasizes that a near-zero value in any one domain can weaken the whole system.
AI literacy and diagnostic expertise are complementary, but must be sequenced
The educational evidence should prevent a one-sided narrative. Well-designed AI training can improve knowledge and performance; inadequate case diversity, feedback, and mentorship create countervailing risks.
AI-focused training can upskill trainees
Median performance before and after AI-focused radiology training (Figure 7).
Source: Keshavarz et al. (2025/2026), Academic Radiology, doi:10.1016/j.acra.2025.10.049. Summarizes 5 of 14 heterogeneous studies; not a pooled effect.

Learning with AI vs. learning from AI
Learning with AI accelerates practice and provides immediate prompts. It produces short-term task support.
Learning from AI requires the trainee to understand why a suggestion is correct, when it fails, what evidence changes the diagnosis, and how to perform when the tool is absent. This builds transferable expertise.
A simulator that always announces a lesion can teach a search strategy that fails in ordinary clinical work.
Figure 8 · A protected learning architecture for AI-assisted interpretation

Author-developed framework informed by Tejani et al. (2023), Kitamura et al. (2025), Keshavarz et al. (2025/2026), and Schramm et al. (2026).
From foundational trainee to clinical leader
| Developmental stage | Primary learning need | Recommended AI sequence | Required evidence of competence |
|---|---|---|---|
| Foundational trainee | Visual search, normal anatomy, pattern library, differential construction, uncertainty awareness. | Substantial protected unaided interpretation; AI revealed after first pass; faculty review of discordance. | Unaided performance, search completeness, confidence calibration, explanation of failure modes. |
| Intermediate trainee | Integrate findings with clinical context, manage ambiguity, learn prioritization. | Mixed workflow by task; independent-first for high-consequence cases; guided AI comparison. | Appropriate use, override reasoning, rare-case exposure, recovery during AI-off simulation. |
| Senior trainee / fellow | Speed without premature closure, consultation, supervision, system stewardship. | Realistic AI-assisted workflow with periodic blinded and AI-off assessment; peer teaching. | Stable combined and unaided performance, escalation judgment, teaching capability. |
| Practicing radiologist | Maintain subspecialty expertise, adapt to tools, detect drift in self and model. | Task-segmented AI use; longitudinal calibration feedback; skill-sensitive audit. | Performance by agreement state, outcome closure, continuing education, downtime readiness. |
| Clinical leader | Govern task design, equity, learning systems, and organizational resilience. | Access to model and human performance dashboards; authority to change sequence or pause use. | Documented decisions, policy compliance, corrective action, evidence of capability preservation. |
Delaying all AI until late training leaves residents unprepared for contemporary practice. Introducing unrestricted AI before foundational reasoning is established can inhibit skill formation. The answer is staged exposure with explicit independent-performance gates.
Five operating domains that make oversight a measurable system
Exposure provides the raw cases. Reasoning converts exposure into hypotheses. Feedback corrects and enriches them. Calibration aligns confidence and reliance with the evidence. Resilience keeps capability available when routine support fails. Governance supplies the owner, policy, data, time, and psychological safety across all five.
Figure 9 · The Radiology Expertise Preservation Model

Author-developed executive operating framework, proposed for local validation. It is not a clinical practice guideline.
| Domain | Design objective | Example controls | Failure signal |
|---|---|---|---|
| Exposure | Maintain variety, rarity, difficulty, normal comparators, and graded autonomy. | Case-mix monitoring; sampled review of AI-negative cases; curated rare-case sets; equitable rotation. | Stable volume with narrowing diversity or declining rare-event contact. |
| Reasoning | Preserve independent search, synthesis, and uncertainty representation when consequences warrant. | Delayed AI reveal; pre-AI findings and confidence; structured differential; protected time. | Minimal independent record; immediate acceptance; high concordant error. |
| Feedback | Connect decisions to outcomes and make discordance educational. | Pathology and follow-up linkage; peer learning; automated discrepancy feeds; coaching. | Repeated error without learning closure; low follow-up completeness. |
| Calibration | Match reliance to model, task, case, and reader conditions. | Agreement-state metrics; override appropriateness; confidence shift; failure-mode simulations. | Overreliance, underreliance, or wide individual variability without intervention. |
| Resilience | Maintain safe operation during outage, drift, novelty, and cyber disruption. | AI-off drills; downtime worklists; manual prioritization; escalation rehearsal; recovery time objective. | Performance collapse, backlog escalation, or unsafe workarounds when AI is absent. |
Which oversight design keeps the human engaged enough to catch failure?
| Oversight mode | Human role | Expertise implication | Appropriate control |
|---|---|---|---|
| Human in the loop | Action required for every case. | Exposure may be preserved, but verification can still become passive. | Independent-reasoning sample, agreement-state monitoring, adequate review time. |
| Human in parallel | Human and AI review independently, then reconcile. | Strongest independent-performance signal; highest labor requirement. | Use selectively for high-consequence validation, training, or drift investigation. |
| Human on the loop | AI acts; humans supervise and intervene. | Rare intervention can weaken recovery and vigilance. | Alert-quality monitoring, simulation, minimum-intervention practice, escalation authority. |
| Human over the loop | Humans govern the system, not individual outputs. | Frontline exposure may disappear; governance expertise becomes critical. | Robust surveillance, audit sampling, multidisciplinary owner, rapid rollback. |
| Human out of the loop | AI performs and acts without contemporaneous review. | Maximum exposure, displacement, and dependency risk. | Restrict to evidence-supported, recoverable tasks with rigorous validation and contingency. |
Risk attaches to the work unit, not to AI as a single category
A single examination contains tasks in different risk categories. Automated segmentation may be highly recoverable; autonomous exclusion of a subtle high-consequence diagnosis may not be. Use consequence and recoverability to place each task.
Task delegation calculator
Set the clinical consequence and recoverability of a specific task to see the recommended delegation mode. This mirrors the executive delegation matrix; example placements are illustrative and require local clinical, regulatory, legal, and technical review.
Figure 10 · Executive AI delegation matrix

Delegation by clinical consequence and recoverability. Example task placements are illustrative; local clinical, regulatory, legal, and technical review is required.
Seven decisions to specify per task
| Task policy element | Required decision |
|---|---|
| Intended use | What exact decision or work product will AI influence, for which population, modality, site, and user? |
| Sequence | Will AI be visible before, during, or after independent interpretation? |
| Human authority | Who may accept, override, escalate, or suspend the output, and what evidence is required? |
| Recovery | How can an error be detected and reversed before patient harm? |
| Learning sensitivity | Does automation remove repetitions, case exposure, or graded autonomy needed for competence? |
| Monitoring | Which model, human, combined, equity, educational, and resilience measures will be reviewed? |
| Stop rule | What threshold triggers investigation, sequence change, retraining, rollback, or suspension? |
The safest workflow is not always the one with the most human steps. It is the one that preserves the right human capability at the point where error is consequential and difficult to recover.
Monitor the model, the human, the team, the learning environment, and AI-off capability
Model monitoring is indispensable but insufficient. It cannot reveal whether clinicians are becoming dependent, whether trainees are losing exposure, or whether combined performance varies across readers. Treat human-AI collaboration as a coupled clinical process.
Measure four performance states
Typical monitoring focuses on AI-alone and combined performance; expertise preservation demands all four.
Author-developed comparison of monitoring focus. Illustrative coverage, not measured values.
The four states
Human alone
What capability exists without the tool? Blinded sample, simulation, delayed reveal, or AI-off drill.
AI alone
What can the model do on local data now? Site validation, silent mode, drift and subgroup monitoring.
Human plus AI
Does the combined workflow improve patient-relevant performance? Prospective comparison, agreement-state review.
Human after interruption
Can the service recover during outage, novelty, or drift? Scheduled downtime simulation and after-action review.
If model sensitivity remains constant while unaided clinician sensitivity declines, the dashboard may stay green until the model is incorrect or unavailable. Monitoring only the algorithm can create a false sense of assurance.
The minimum analytic unit is the four-state agreement grid
Human correct / AI correct
Is speed gained without weakening independent search? Sample for premature closure and opportunities to simplify low-risk workflow.
Human correct / AI wrong
Did the human resist the error, and what evidence supported the override? Reinforce appropriate override; investigate model failure.
Human wrong / AI correct
Was useful AI information seen, understood, and acted on? Improve presentation, training, escalation, or confidence communication.
Human wrong / AI wrong
Do humans and models share a blind spot? Escalate to failure-mode analysis, education, workflow redesign, or additional safeguards.
What to review, and how often
| Domain | Measure | Review cadence | Guardrail question |
|---|---|---|---|
| Quality | Combined diagnostic performance (sensitivity, specificity, critical miss rate). | Monthly | Is combined performance improving for all material subgroups? |
| Independent capability | Unaided performance on blinded, delayed-reveal, or AI-off cases. | Quarterly | Is human-alone capability stable within tolerance? |
| Reliance | Override appropriateness after adjudication. | Monthly → quarterly | Are users over-relying, under-relying, or varying unpredictably? |
| Calibration | Confidence shift by agreement state. | Quarterly | Does AI increase confidence even when wrong? |
| Exposure | Learning exposure index by rarity, difficulty, and autonomy. | Monthly (trainees) | Which cases have disappeared from human practice? |
| Feedback | Outcome closure rate. | Monthly | Are decisions converted into learning? |
| Resilience | AI-off recovery performance. | Semiannual drill | Can the service sustain a safe degraded mode? |
| Workforce | Cognitive burden and trust calibration. | Baseline, 30/60/90 days | Is safety achieved through unsustainable hidden work? |
| Equity | Benefit and harm distribution by site, shift, and role. | Quarterly | Who gains, who bears verification work, who loses learning? |
Clinician-level data are necessary because AI effects vary by user, and they are sensitive. If dashboards rank productivity or punish disagreement, users conceal uncertainty and game the workflow. Default use should be system improvement and confidential development, with escalation reserved for validated, repeated safety concerns.
Expertise risk crosses many domains, so ownership must be explicit
A committee can coordinate evaluation, cybersecurity, privacy, procurement, equity, and monitoring, but it does not substitute for an accountable executive. Each use case needs a named clinical owner with authority to change sequence, require education, commission an audit, modify a threshold, suspend use, or initiate rollback.
Workflow sequence can carry legal meaning
Mock-juror responses in a hypothetical missed-bleed case, by radiologist-AI sequence (Figure 11).
Source: Bernstein et al. (2026), Nature Health, doi:10.1038/s44360-026-00085-2. Randomized study of 282 mock jurors. Behavioral research, not legal advice.

When AI correctly flagged the CT, 74.7% of mock jurors sided with the plaintiff if the radiologist read once after seeing AI, versus 52.9% when the radiologist interpreted independently first. The odds of siding with the plaintiff were 2.6 times higher in the single-read condition.
Independent-first documentation is not a liability shield and may be inappropriate for some tasks. Obtain jurisdiction-specific legal review. The broader point: workflow architecture, training, alerts, and documentation can become part of the factual record when harm occurs.
Nondelegable accountability by role
| Role | Nondelegable accountability |
|---|---|
| Board / quality committee | Confirm that AI risk oversight includes human capability, resilience, and equity, not only vendor and cybersecurity risk. |
| Chief medical / clinical officer | Set enterprise clinical AI principles, resolve cross-service risk, and ensure accountability for patient outcomes. |
| Radiology chair / executive | Own task policy, staffing and education implications, performance review, and clinical escalation. |
| Program director / education leader | Protect graded autonomy, exposure portfolio, AI literacy, and independent competency assessment. |
| AI governance committee | Standardize intake, evidence review, validation, monitoring, change control, and incident response. |
| Quality & safety leader | Integrate AI agreement states and expertise signals into peer learning, root cause analysis, and safety reporting. |
| Vendor / technology partner | Provide performance evidence, intended-use limits, version transparency, logging, update notice, and audit/rollback support. |
| Frontline user | Use within policy, document material disagreement, report failure, participate in training, and retain authority to escalate. |
From tool adoption to an expertise-resilient system
Level 1
Tool adoption. Model integrated; success defined by accuracy or speed. Expertise risk is assumed to be controlled by having a human reviewer.
Level 2
Clinical validation. Local performance and intended use evaluated; users trained. Automation bias recognized, but unaided capability and exposure are not measured.
Level 3
Coupled-system governance. Human, model, and combined performance monitored by agreement state and subgroup. Task policies and learning safeguards are active.
Level 4
Expertise-resilient system. Longitudinal capability, education, AI-off recovery, equity, and change control integrated. Expertise treated as a strategic safety asset.
Test one consequential use case, learn quickly, create a reusable standard
The first pilot should be clinically meaningful but bounded: an existing or near-term deployment, measurable outputs, accessible outcome data, a defined user group, and a plausible expertise-sensitive mechanism.
Figure 12 · A 90-day pilot from baseline to scale decision

Author-developed implementation roadmap. Timing should be adjusted for use-case complexity, regulatory obligations, data readiness, and local change processes.
Six phases, each with analytic outputs
Name the owner; define the frame
Intended use, sequence, users, exclusions, escalation, stop rules, and data governance.
Outputs: signed use-case charter; risk register; metric dictionary; training plan.
Measure the starting point
Current performance, case mix, confidence where feasible, feedback closure, workload, and downtime process.
Outputs: baseline dashboard; data-quality assessment; power and sampling plan.
Train and launch a protected trial
Train on model strengths, failure modes, overrides, reporting, and an independent-first workflow.
Outputs: competency evidence; technical validation; early burden review.
Review discordance weekly
Correct integration and interface failures; monitor safety and user burden.
Outputs: agreement-state analysis; incident log; process fidelity; corrective actions.
Run rare-case and AI-off exercises
Assess recovery, escalation, and independent performance under simulated absence.
Outputs: simulation results; downtime performance; learning gaps; remediation plan.
Compare to baseline and guardrails
Evaluate subgroup, educational, legal, and operational implications.
Outputs: scale / revise / pause / stop recommendation; reusable policy and lessons learned.
How expertise-resilient is your current deployment?
Check each control your service line has in place. The meter maps to the report’s four-level maturity model.
Named accountable clinical owner
One executive has authority to change sequence, pause use, or initiate rollback.
Task-level delegation policy
Each task is placed by consequence and recoverability, not labeled broadly as “AI-assisted.”
Independent-first workflow where warranted
Protected cognition captured before AI reveal for high-consequence, educationally sensitive tasks.
Unaided performance monitored
Human-alone capability tracked via blinded, delayed-reveal, or AI-off samples.
Agreement-state review
Cases classified into the four agreement states; overrides adjudicated for appropriateness.
Learning exposure protected
Case-mix quotas, rare-event simulation, and graded autonomy for trainees.
AI-off resilience drills
Scheduled downtime exercises with measured recovery, escalation, and backlog performance.
Board-visible guardrails and stop rules
Prespecified thresholds trigger investigation, sequence change, retraining, or rollback.
A workflow that improves combined sensitivity but materially worsens unaided performance, trainee exposure, or cognitive burden presents a leadership tradeoff, not an automatic success. Keep outcomes visible as a balanced set and document the decision rationale.
Selected literature and guidance
References are current through July 17, 2026. Filter by evidence stream; each links to the DOI, publisher, or authoritative source. This report is a rapid executive synthesis, not a registered systematic review.
Performing best when needed least: Reader experience shapes accuracy gains in LLM-assisted brain MRI differential diagnosis. Radiology, 319(2).
Automation bias in mammography: The impact of AI BI-RADS suggestions on reader performance. Radiology, 307(4), e222176.
Heterogeneity and predictors of the effects of AI assistance on radiologists. Nature Medicine, 30, 837-849.
Endoscopist deskilling risk after exposure to AI in colonoscopy: A multicentre, observational study. Lancet Gastroenterol Hepatol, 10(10), 896-903.
AI in medicine: A scoping review of the risk of deskilling and loss of expertise among physicians. ESMO Real World Data & Digital Oncology, 12, 100693.
Does using AI assistance accelerate skill decay and hinder skill development without the performers’ awareness? Cognitive Research, 9, 46.
AI-induced deskilling in medicine: A mixed-method review and research agenda. Artificial Intelligence Review, 58, 356.
AI education in radiology training: A systematic review of effectiveness, barriers, and future directions. Academic Radiology, 33(3), 695-706.
Teaching AI for radiology applications: A multisociety-recommended syllabus from the AAPM, ACR, RSNA, and SIIM. Radiology: AI, 7(6), e250137.
Artificial intelligence and radiology education. Radiology: Artificial Intelligence, 5(1), e220084.
AI for radiology: A primer, part II. Interacting with AI results. Radiology, 320(1).
Developing, purchasing, implementing, and monitoring AI tools in radiology: A multisociety statement (ACR, CAR, ESR, RANZCR, RSNA). Radiology: AI, 6(1), e230513.
Postdeployment monitoring and surveillance methods, guidelines, and possibilities for AI in radiology. RadioGraphics, 46(7), e250173.
Alignment of policy, practice, and patient safety for trustworthy AI in radiology. Radiology: Artificial Intelligence.
FUTURE-AI: International consensus guideline for trustworthy and deployable AI in healthcare. BMJ, 388, e081554.
Guidance on the responsible use of artificial intelligence in healthcare.
The radiologist-AI workflow and the risk of medical malpractice claims. Nature Health, 1, 386-389.
Understanding liability risk from using healthcare AI tools. New England Journal of Medicine, 390(3), 271-278.
Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381-410.
Automation bias: A systematic review of frequency, effect mediators, and mitigators. JAMIA, 19(1), 121-127.
The false hope of current approaches to explainable AI in health care. Lancet Digital Health, 3(11), e745-e750.
The potential expertise paradox in AI-assisted radiology. Radiology, 319(2).