Reach out to me for a complete copy of the research report

Executive Research Report · Interactive Edition
AI Will Not Fix Healthcare’s Broken Workflows. It Will Industrialize Them.
A workflow-first leadership model for safe, accountable, and value-producing healthcare AI. Every figure, table, gate, and index from the research report, rebuilt as working tools.
- Kelly Emrick, DHSc, PhD, MBA, BSRT(ARRT)R
- E-R-A-A sequence navigator
- AWRI readiness index
- Evidence current through August 20, 2026
- 18 sections
Abstract and hypothesis
A provocative leadership claim, tested against the implementation evidence
Healthcare organizations are investing rapidly in artificial intelligence while many of the processes selected for automation remain poorly mapped, weakly owned, and burdened by accumulated workarounds.
Hypothesis verdict
Supported as a leadership and implementation proposition, with moderate confidence.
Evidence convergence is strong across sociotechnical theory, systematic reviews, and real-world cases. Direct head-to-head causal evidence comparing sequencing strategies remains limited, so the title should be read as a disciplined warning, not a claim that every AI deployment will fail.
Board-level implication
Approve AI as a change to the operating model, not as a software purchase. No production deployment should proceed without an agreed process owner, clinical owner, safety owner, measurable patient outcome, total-work metric, exception pathway, and stop rule.
This report tests the hypothesis that when AI is introduced before low-value work is eliminated, the end-to-end workflow is redesigned, human accountability is explicit, and the process is validated, it is more likely to industrialize dysfunction than to produce durable clinical and operational value.
The synthesis integrates systematic reviews, real-world implementation studies, human factors experiments, implementation science, clinical governance literature, and authoritative health system guidance available through August 20, 2026.
The evidence is directionally supportive but does not prove a universal causal law. Direct comparative trials of workflow-first versus product-first AI implementation are scarce. Convergent evidence does show that model accuracy alone does not predict clinical value, that apparent time savings often fail to demonstrate lower system workload, and that automation can propagate biased proxies, note bloat, alert fatigue, and new verification work.
Four research questions, four executive answers
| Research question | Executive answer |
|---|---|
| RQ1. Does AI create value independently of workflow design? | No. Evidence consistently treats value as an interaction among technology, work system, users, accountability, and context. |
| RQ2. Can automation reduce one task while increasing total work? | Yes. Review and field evidence documents verification, exception handling, longer outputs, alerts, and downstream work that task-time measures can miss. |
| RQ3. What sequence should leaders require? | Eliminate low-value work, redesign the entire pathway, assign accountability, validate the process, then automate. |
| RQ4. What should boards govern? | Patient outcomes, total system work, human reliability, equity, incident response, drift, and lifecycle ownership, not adoption counts or model accuracy alone. |
What the evidence base looks like
meta-analyses found a statistically significant task-time effect, even though 67% of task-time studies reported reductions.
of 25 reviewed AI implementation frameworks covered Act, the domain where post-deployment learning happens. Plan reached 84%.
of internal and emergency medicine physicians in an experiment followed every piece of incorrect advice they were given.
of 77 healthcare AI governance frameworks covered principles, assessment, lifecycle, and oversight together.
What this dashboard adds to the report
Every number below is computed live in your browser from the report’s own published figures and formulas. Nothing is sent anywhere.
Five working engines
The E-R-A-A Sequence Navigator, the AWRI readiness calculator, the Industrialized Dysfunction Risk screener, the Net Work Ledger, and the alert burden model each implement a rule from the report rather than describing it.
Rules enforced, not displayed
The navigator refuses to advance past an open gate. The AWRI reports a stop rather than a lower score when a noncompensable gate fails. The dysfunction screener reports both verdicts when its two decision rules disagree.
Three findings derived here
The AWRI band table and its own stop gates disagree; the dysfunction model’s product ranking and its two-high rule disagree; and the report’s alert-burden arithmetic can be reconstructed from the published rates.
Where the derived findings sit
A score of 91 that still cannot ship
Of the 28,901,376 whole-point ways to score the six AWRI domains, 229,390 land in the top Scale consideration band. 39,088 of those, 17.04%, fail at least one stop gate. The highest total that can fail the accountability gate is 91 of 100. Open the calculator.
Two rules that rank differently
Scored 1 to 5, the dysfunction model’s four domains produce 625 combinations. Its multiplicative product and its published two-high rule disagree on 174 of them, 27.84%. A workflow at 3-3-3-3 scores 81 and is never flagged; 5-5-1-1 scores 25 and always is. Open the screener.
The alert burden behind AUC 0.63
The published external validation rates rebuild to roughly 6,922 alerts producing about 831 true positives, 8.3 alerts per true positive, with an implied 2,517 sepsis cases and about 1,686 missed. Open the model.
All three are author-derived from the report’s published values and are labeled as such wherever they appear. They are arithmetic consequences of the report’s own numbers, not new empirical claims.
Central thesis
The claim in one paragraph
Artificial intelligence amplifies the operating system it enters. When deployed in an unexamined workflow, it can increase the speed, scale, and consistency of waste, rework, alert burden, inequity, and accountability failures, and increase their opacity.
The proper unit of transformation is the care process, not the model.
The report this dashboard is built from

The report runs to nine figures, twenty-nine tables, three original frameworks, a twelve-study evidence matrix, a twelve-metric dictionary, and a twenty-item pre-automation checklist. Every one of those elements appears in this dashboard, either as the published figure, as an interactive rebuild, or as a working tool.
Where a figure prints its own values clearly, it is shown as published rather than rebuilt. Where an interaction, a derived series, or legibility justifies a rebuild, the rebuild sits alongside a toggle that reveals the published original, so the two can always be compared.
The three original frameworks are labeled throughout as proposals requiring local testing. The report is explicit that they are not validated instruments, and this dashboard repeats that caveat at the point of use rather than burying it in a note.
Keywords
artificial intelligencehealthcare leadershipworkflow redesignsociotechnical systemshuman factorsimplementation scienceclinical governancepatient safetyworkforce well-beinghealth equity
The industrialization thesis
Automation changes the production characteristics of a workflow before it changes its purpose
Industrialization is used here in a precise operating sense. AI can make a process faster, more repeatable, cheaper per transaction, more scalable, and less dependent on individual effort. Those properties are valuable only when the process produces the right result. If the work is unnecessary, the decision target is wrong, exceptions are unsafe, or responsibility is diffuse, industrialization widens the defect’s reach.
The unit of analysis, therefore, cannot be the model alone. Sittig and Singh’s sociotechnical model describes health information technology as interacting with hardware and software, clinical content, people, workflow and communication, organizational policy, external rules, and measurement. The SEIPS tradition similarly links work-system design to process, patient, professional, and organizational outcomes. NASSS adds the condition, technology, value proposition, adopters, organization, wider context, and adaptation over time.
Across these frameworks, technology is one component of a complex adaptive system, not an independent treatment.

The same six capabilities, two different outcomes
Each AI capability listed below is neutral. What determines its effect is the condition of the workflow it enters. Use the switch to see the report’s paired reading of the same capability.
| AI capability | When the workflow is sound | When the workflow is broken |
|---|---|---|
| Speed | Shortens time to a necessary action | Accelerates unnecessary work and premature decisions |
| Scale | Extends reliable service to more patients | Expands inequity, alert burden, and low-value activity |
| Consistency | Reduces unwarranted variation | Standardizes a flawed rule or proxy |
| Generativity | Produces useful drafts and summaries | Creates note bloat, review burden, and confidence errors |
| Prediction | Prioritizes a response with defined capacity | Creates queues and alerts without response resources |
| Opacity | May be manageable with controls and monitoring | Makes accountability and failure detection harder |
The critical distinction
A technically successful deployment can be an operational failure.
High uptime, rapid inference, and widespread adoption do not translate into improved patient outcomes, reduced total workload, safer decisions, or greater equity.
The two implementation sequences
The report’s first figure sets up the whole argument: the same organization, the same technology, two different orders of operations. The interactive rebuild lets you step through either track and see what each step assumes.
Product-first and workflow-first sequences, stepped
Select a track, then step through it. The dashed return line on the product-first track is the report’s point: the result feeds back into the tool decision rather than into the process.
Figure rebuild. Interactive version of the report’s Figure 1. Source: Original author framework. The sequence is normative and has not been validated as a comparative instrument.

How the evidence was weighted
The report used a structured critical synthesis rather than a de novo systematic review. Priority was given to peer-reviewed systematic and scoping reviews, real-world implementation studies, randomized and quasi-experimental evaluations, human factors experiments, consensus guidelines, and foundational sociotechnical frameworks. The 2026 American Hospital Association workforce scan was treated as industry context, not equivalent to peer-reviewed evidence. The 2025 JAMA Summit report was treated as a multidisciplinary consensus synthesis.
| Evidence tier | Examples in this report | Permitted inference |
|---|---|---|
| Tier 1: comparative clinical evidence | Randomized trials, quasi-experimental evaluations, and controlled implementation studies | Association or causation only as supported by design; context and implementation bundle remain explicit |
| Tier 2: evidence synthesis | Systematic reviews, scoping reviews, reviews of implementation frameworks | Convergence, gaps, heterogeneity, and recurring determinants |
| Tier 3: human factors and field studies | Workflow observation, simulation, user studies, usability research | Mechanisms, failure modes, and design requirements |
| Tier 4: consensus and authoritative guidance | FUTURE-AI, DECIDE-AI, JAMA Summit, AHA scan | Normative governance expectations and practice context |
| Tier 5: original proposals | E-R-A-A, Industrialized Dysfunction Risk, AI Workflow Readiness Index | Decision aids requiring local testing; not presented as validated instruments |
Interpretive rule
Measurement, inference, and proposal are separated. Empirical figures reproduce reported study results. Original equations and indices are clearly labeled conceptually. Causal language is reserved for designs that support it.
Limitations of the method include the absence of protocol registration, dual independent screening, formal risk-of-bias scoring for every included source, and quantitative meta-analysis. The report is suitable as an executive research synthesis and as a journal-ready foundation, but a target journal may require conversion into a formal systematic, scoping, or realist review.
Healthcare already has workflow debt
AI does not enter a clean production system
It enters layered policy, technology, and workarounds.
Workflow debt is the accumulated burden that arises when new requirements, fields, approvals, alerts, handoffs, and workarounds are added without removing obsolete work or redesigning the entire pathway. Like technical debt, it may be locally rational and globally damaging. Each workaround keeps care moving today while making tomorrow’s process harder to understand, measure, and govern.
The electronic health record makes the debt visible
Where the clinical day goes
Three independent measurements of the same underlying problem, drawn from the report’s cited time-motion and event-log evidence.
Chart source. Values as reported in the research report from Sinsky et al. (2016), Arndt et al. (2017), and more recent national evidence. The 51.8% share is arithmetic from the reported 5.9 of 11.4 hours and is labeled as derived.
Time-motion research found that for every hour of direct clinical face time, physicians spent nearly two additional hours on electronic records and desk work. Event-log analysis later found that primary care physicians spent nearly 6 hours of an 11.4-hour workday in the record, including substantial after-hours work. More recent national evidence continues to show that electronic work consumes nearly half of clinic time.
These burdens are not caused by software alone. Documentation rules, billing incentives, inbox routing, staffing models, and local policy shape the work.
Why this matters for AI
An organization that automates documentation without changing documentation requirements has bought a faster way to satisfy the requirement. The requirement is the debt.
Six debt patterns, and why each resists direct automation
| Common debt pattern | Observable signal | Why direct automation is risky |
|---|---|---|
| Duplicate documentation | The same fact is entered, copied, reconciled, and restated | Generation can increase volume without reducing the underlying requirement |
| Approval accumulation | Multiple signatures or authorizations with unclear risk logic | Automation can conceal unnecessary control points and make appeals harder |
| Alert accumulation | High override rates, low action yield, repeated warnings | Prediction increases queue volume unless thresholds and response capacity are redesigned |
| Inbox fragmentation | Messages bounce among clinicians, nurses, pools, and schedulers | Drafting a reply does not fix routing, scope, ownership, or demand |
| Workarounds | Shadow lists, manual tracking, copy-forward, verbal escalation | AI may learn the workaround as if it were the intended process |
| Unowned exceptions | A normal path exists, but failure states lack an accountable owner | Automation handles the easy majority while concentrating risk in the difficult minority |
Leadership test
If leaders cannot draw the current-state pathway, name its purpose, quantify its failure demand, identify its exceptions, and assign ownership, they are not ready to automate it at scale.
Score your own debt exposure
Mark the patterns you can observe in the workflow you are considering for automation. This is a structured reading of the table above, not a validated instrument.
Workflow debt exposure check
Six patterns. Each one you can observe is one the report says automation will amplify rather than resolve.
Debt patterns present
0 of 6
Mark the patterns you can observe today.
Scoring rule used here: this is a count, not a weighted index. The report does not assign weights to these six patterns, so none are invented. The unowned-exception pattern is reported separately because the report treats it as the pattern that concentrates risk rather than spreading it.
AI often enters too late in improvement
The product decision should follow process design, not substitute for it.
Many organizations begin with a capability inventory: summarization, prediction, computer vision, coding, scheduling, or agentic execution. They then seek use cases that justify a product already selected. This sequence imposes a design constraint on the current workflow. The organization asks where the technology fits rather than whether the work should exist, who should do it, and what result the patient needs.
Implementation reviews repeatedly identify workflow fit, stakeholder engagement, local adaptation, infrastructure, reflection, and evaluation as determinants of uptake. Preti and colleagues reviewed 34 empirical implementations and found recurring importance of design, relative advantage, complexity, access to knowledge, infrastructure, stakeholder engagement, and evaluation. Nair and colleagues combined 38 empirical cases with 69 stakeholder interviews, organizing barriers and strategies across leadership, change management, engagement, workflow, resources, legal issues, training, data, monitoring, maintenance, and ethics.
Read this correctly
These are not peripheral adoption concerns. They define the intervention.
Reversing the buying sequence
Six questions that decide whether you are running product-first or workflow-first
The buying sequence is often reversed: a product is selected, use cases are solicited, and the current workflow becomes an immutable constraint. This converts transformation into tool insertion.
The report’s diagnostic is unusually practical. You do not need a maturity survey to find out which sequence your organization is running. You only need to listen to the question being asked in the room. Each product-first question below has a workflow-first replacement that answers something the first question cannot.
Question translator
Select the question your organization is actually asking about a candidate AI workflow. The translator returns the workflow-first replacement, the E-R-A-A stage the replacement belongs to, and the evidence that answers it.
Which question is being asked in your room right now?
Product-first question
Where can we deploy this capability?
This question assumes the capability is already chosen, so the workflow becomes a place to put it rather than the thing being improved.
Workflow-first replacement
Which patient or workforce outcome is currently constrained, and why?
Where this question belongs in the sequence
Stage 2
Redesign the end-to-end workflow
The full replacement table
| Product-first question | Workflow-first replacement |
|---|---|
| Where can we deploy this capability? | Which patient or workforce outcome is currently constrained, and why? |
| How many users can we activate? | How much low-value work can we remove before deployment? |
| What is model accuracy? | What is the reliability of the human-plus-AI pathway under normal and exceptional conditions? |
| How much task time is saved? | How much total work is removed, shifted, created, or released to care? |
| Did clinicians accept the tool? | Did the process improve outcomes without worsening safety, equity, or workload? |
| Can the vendor monitor it? | Who in the health system has authority to change, pause, or retire it? |
Procurement principle
A request for proposal should describe the validated problem, target outcome, workflow requirements, accountable roles, data boundaries, monitoring obligations, and exit conditions before it describes preferred AI functionality.
What the replacement costs you
It slows the first decision
Workflow-first questions cannot be answered by a vendor deck. They require a current-state map, a baseline, and a named owner, which takes weeks that a product-first process does not spend.
It moves the risk earlier
The report’s sequencing rule is that skipping a stage does not save time. It transfers time and risk into production, where defects are more expensive, patients are exposed, and organizational learning becomes harder.
It should not become perfectionism
Processes in complex care will never be fully stable. The disciplined alternative is bounded uncertainty: define the problem, remove obvious waste, design responsibility, test on live local data without influence, deploy reversibly, measure total effects, and retain authority to stop.
The efficiency evidence is more fragile than headlines suggest
Savings measured within a single task may disappear across the entire work system
Efficiency is one of the most common executive justifications for healthcare AI, but its measurement is frequently narrow. A model may shorten interpretation, drafting, or review while creating new work in verification, exception handling, patient communication, downstream testing, appeals, training, maintenance, or incident response. If the saved minutes cannot be translated into released capacity, lower cost, better access, or improved care, the result is local acceleration rather than system value.
Wenderott and colleagues screened 13,756 records and included 48 original studies of AI in routine medical imaging. Of 33 studies measuring task time, 67% reported reductions. Yet three meta-analyses using 12 studies found no statistically significant effect, and only three studies closely examined how time saved translated into workload. The review identified multiple workflow patterns, underscoring that the effect of AI depends on where and how it changes work.
The evidence funnel, with the denominators the headline drops
Each stage is drawn to the count reported in the review. The panel to the right of each stage shows what share of the previous stage survived.
Chart source. Wenderott et al. (2024) as reported in the research report. The stage-to-stage retention percentages and the count of studies reporting reductions are arithmetic from those published figures and are labeled as derived.

What the funnel arithmetic shows
48 of 13,756 records became included studies. The base is narrow before any effect is estimated.
67% of 33 resolves to 22 studies, since 22 of 33 is 66.67% and 23 of 33 is 69.7%. Derived here.
3 of the 48 included studies scrutinized whether saved time reduced workload. Derived here.
meta-analyses, using 12 studies, found a statistically significant effect.
The gap the headline conceals
Twenty-two studies reporting a reduction and zero meta-analyses finding a significant pooled effect are not contradictory. They are two different questions. The first asks whether a task got shorter in a study. The second asks whether the evidence, combined, supports the claim. Only the second supports an investment case.
Five measurement levels, and what each one can miss
The report’s measurement ladder is the reason a task-time claim cannot stand alone. Select a level to see what it captures and what it cannot see.
What it measures
Model
Accuracy, AUROC, sensitivity, calibration
What it can miss
Blind spots at this level
Response capacity, workflow fit, action, patient outcome
| Measurement level | Example metric | What it can miss |
|---|---|---|
| Model | Accuracy, AUROC, sensitivity, calibration | Response capacity, workflow fit, action, patient outcome |
| Task | Seconds per image, note, message, or chart | Review, exceptions, corrections, handoffs, downstream demand |
| Role | Clinician time or perceived burden | Work shifted to nurses, staff, patients, or caregivers |
| Pathway | Episode cycle time, closure, avoidable delay | External costs, equity effects, long-term maintenance |
| Enterprise | Capacity released, cost, access, safety, workforce retention | May require longer observation and causal attribution |
Minimum efficiency claim
Do not claim efficiency from reduced task time alone.
Demonstrate net change in total work per completed episode, including verification, exceptions, rework, maintenance, and downstream utilization, then show how released capacity was actually used. The Net Work Ledger in the next section implements exactly this rule.
Total work, not task time
The Net Work Ledger
The report sets a minimum standard for any efficiency claim. This ledger is that standard, made arithmetic.
An efficiency claim is permitted only when net total work per completed episode falls, and only when the released capacity is shown to have gone somewhere. The ledger below takes the task-level saving an organization is being sold, adds the work the report says accompanies it, and reports both numbers side by side so the difference is visible rather than assumed.
Net work per completed episode
All inputs are yours. Nothing is prefilled from any organization’s data. The four loaded profiles use the published effect sizes from the report’s cited studies for the task saving only; every other line is an explicit local assumption you should replace.
Baseline and saving
An episode is one completed unit of the work: a visit documented, a message resolved, a study read.
All roles, from demand to resolved outcome. Not just the clinician.
The number the efficiency claim is based on, as reported by the vendor or the pilot.
Work the claim does not count
Longer records to read, extra tests, patient calls, appeals.
Lifecycle operations. The report treats this as recurring clinical infrastructure, not a project cost.
The report is explicit that estimated vendor hours do not count. Only observed capacity created and intentionally redeployed does.
Who absorbs the added work
Shares are normalized to 100%. This does not change net work. It changes who carries it, which the report treats as a separate question of equal weight.
Net change in total work per episode
+0.41 min
Total work rises while task time falls
Where the added work lands
Monthly hours of added work by group, after normalization. Where net work falls, this chart shows the distribution of relief instead, and the label changes.
The ledger
| Line | Minutes per episode | Hours per month | Counted by a task-time claim? |
|---|
Waterfall of the ledger above. A bar below the axis removes work, a bar above it adds work, and the final column is the net position the report asks leaders to report instead of the task saving.
What this ledger cannot tell you
A favorable net position is a necessary condition for an efficiency claim, not a sufficient one.
The report requires the same deployment to show that patient outcomes did not worsen, that safety and equity held, and that the released capacity improved care rather than only increasing throughput. Those are separate measurements with separate owners, and none of them is arithmetic.
Planning without learning produces brittle implementation
Organizations are better at approving AI than at adapting it
A systematic review of 25 AI implementation frameworks mapped their content to the Plan-Do-Study-Act framework. Planning appeared in 84% of frameworks, study in 60%, doing in 52%, and acting in only 24%.
The pattern matters. Healthcare AI changes as data, clinical practice, interfaces, staffing, thresholds, and user behavior change. A framework that ends at go-live treats deployment as a finish line when it is actually the beginning of evidence generation.
Framework coverage by PDSA domain
Share of the 25 reviewed frameworks covering each domain, with the underlying framework count shown on each bar.
Chart source. Khan et al. (2024) as reported in the research report. Percentages are published; the framework counts are arithmetic on the denominator of 25 and are labeled as derived. All four resolve to whole frameworks, which confirms 25 as the denominator for every domain.

percentage points between Plan at 84% and Act at 24%. Plan is covered 3.5 times as often as Act.
Derived here from the published 24%. Six frameworks address the phase where a deployed tool is changed, retrained, or retired.
Leadership implication
Frameworks are strongest before deployment and weakest where post-deployment learning must occur. An organization that adopts a framework inherits that shape unless it deliberately funds the Act phase.
Geography of the guidance itself
The same review found geographic inequity in authorship: only one of 172 authors was from a low-income country, and two were from lower-middle-income countries. Frameworks built primarily in well-resourced settings may understate constraints related to infrastructure, staffing, data completeness, maintenance, and access.
The challenge is not only to develop more guidance but to test whether governance is usable in the settings where AI will operate.
0.58% of the authorship of the reviewed implementation frameworks.
1.16%. Together with the line at left, 1.74% of authors.
of the papers in a systematic review of 26 hospital AI implementation studies, which also identified 28 enablers and 18 barriers.
The lifecycle the frameworks stop short of
Where implementation frameworks weaken, the report substitutes a lifecycle with an evidence question and a required control at every stage. Select a stage.
Evidence question
Predeployment
Does the process work without AI, and is the target valid?
Required control
What must be in place
Current-state map, necessity test, baseline outcomes, equity analysis
| Lifecycle stage | Evidence question | Required control |
|---|---|---|
| Predeployment | Does the process work without AI, and is the target valid? | Current-state map, necessity test, baseline outcomes, equity analysis |
| Silent mode | How does the model behave on live local data without influencing care? | Calibration, data quality, subgroup performance, failure review |
| Bounded deployment | Can users safely integrate the output under a real workload? | Training, usability, escalation, downtime, observation |
| Scale | Do benefits persist across sites and populations? | Site-level controls, capacity checks, learning cadence |
| Sustainment | Has drift, work migration, or clinical context changed value? | Outcome monitoring, audit, retraining, and retirement criteria |
Alerts at scale: sepsis as a sociotechnical case
The same clinical intent can produce very different operating consequences
Sepsis prediction illustrates why algorithm evaluation and implementation evaluation cannot be separated. In an external validation of a widely implemented sepsis model, Wong and colleagues analyzed 38,455 hospitalizations. Hospitalization-level AUROC was 0.63. At a threshold of six or higher, the model alerted on 18% of hospitalizations, failed to identify 67% of sepsis cases, and had a positive predictive value of 12%. The burden of alert evaluation was therefore part of the clinical effect.
A later before-and-after quasi-experimental study of COMPOSER involved 6,217 adult patients with sepsis across two emergency departments. Deployment was associated with a 1.9 percentage-point absolute reduction in in-hospital mortality, a 5.0 percentage-point absolute increase in bundle compliance, and lower organ dysfunction. The surrounding implementation included a nurse-facing alert, standardized response options, nurse-to-physician communication, stakeholder engagement, and continuous monitoring of data and performance. Only 5.9% of alerts were dismissed during the intervention period. Benefits were not uniform between the two hospitals, reinforcing the role of context.
Two studies, two different questions
Left: what the model does on its own. Right: what a model plus a redesigned response pathway did. Different designs, populations, endpoints, and settings. These panels are not a comparison of algorithm quality.
Chart source. Wong et al. (2021); Boussina et al. (2024); Kwong et al. (2024), as reported in the research report. Published values only.

Alert burden model
The report argues that alert evaluation is part of the clinical effect, not an operational footnote. The defaults reproduce the published external validation exactly; move any input to see what a threshold decision does to the work it creates.
The report does not publish this value. It is your local assumption and is labeled as such in the readout.
Set to 33% by default because the published figure is that 67% of sepsis cases were missed.
Every 100 alerts, split into the ones that find a case and the ones that do not, at your current positive predictive value.
Alerts per true positive
8.3
alerts evaluated for each case the model finds
How the implied case count is derived
Cases equal alerts multiplied by positive predictive value, divided by sensitivity. At the published rates that is 38,455 x 18% x 12% / 33%, giving about 2,517 cases and a 6.55% prevalence. Because the three published rates are each rounded to whole percentage points, the implied prevalence is only determined to a range of roughly 6.01% to 7.12%. The derivation is author-derived arithmetic on the report’s published values, not a figure the source study states.
What the integrated implementation added
reported alongside a 17% relative change, which implies a baseline in-hospital mortality near 11.2%. Derived here; the rounding range is 10.6% to 11.8%.
reported alongside a 10% relative change, which implies a baseline compliance of exactly 50.0%. Derived here; the rounding range is 47.1% to 53.2%.
during the intervention period, against the alert-evaluation burden modeled above for an unintegrated deployment.
The two implied baselines are arithmetic consequences of reporting the same effect in both absolute and relative terms. They reconcile the report’s text with its Figure 5 and are labeled author-derived. They are not values the source studies publish, and they are sensitive to the rounding of the relative figures.
The intervention is not the model
| Implementation component | Why it matters |
|---|---|
| Right recipient | An alert sent to a person without authority, time, or scope is not an intervention |
| Defined response | Standardized options reduce ambiguity and create measurable process states |
| Communication pathway | Prediction has value only if it accelerates appropriate clinical coordination |
| Capacity | High alert volume without treatment, diagnostic, or escalation capacity creates queues |
| Local validation | Thresholds, prevalence, documentation, and practice patterns change performance |
| Continuous monitoring | Data drift, model drift, and behavior drift can erase the initial value |
Leadership inference
An AI alert should be governed as a redesigned care pathway with a predictive component.
The intervention is the model plus recipient, timing, interface, response standard, resources, accountability, and feedback loop.
Generative AI: relief, expansion, and cognitive debt
The fastest-growing use cases show why output volume is not equivalent to value
Generative AI can relieve a meaningful burden, particularly when it converts an encounter into a draft note or provides a starting point for repetitive communication. It can also increase the quantity of text that humans must verify. The result may be lower creation effort but higher reading, correction, and information-retrieval burden.
This is a form of cognitive debt: future clinicians inherit more fluent material whose accuracy, relevance, and provenance must still be judged.
Reported effects, sorted by direction rather than by study
Relative change from baseline or comparison. Bars to the left reduce time. Bars to the right add time or output. The hatched bar was not statistically significant.
Chart source. Duggan et al. (2025) and Tai-Seale et al. (2024) as reported in the research report. Separate studies and designs; estimates are not directly comparable across interventions. Published values only.

The studies behind the bars
Ambient documentation, 46 clinicians
In a single-group pre-post study of ambient scribing, time in notes per appointment decreased by 20.4% and after-hours work time decreased by 30.0%, while note length increased by 20.6%. Clinicians reported mixed views about note quality and specialty fit.
Principal limitation as stated in the report: short, opt-in, single-system study.
AI-drafted patient replies
In a randomized waiting-list quality improvement study, access was associated with 21.8% greater read time, no significant reduction in reply time, and 17.9% longer replies. Another five-week implementation among 162 clinicians showed a 20% draft utilization rate and improvements in task-load and work-exhaustion measures, but no improvement in time.
Principal limitation as stated in the report: one early tool in an academic system.
Reading these together
Perceived burden and measured time moved in different directions in more than one study. A workforce can feel relief that a ledger does not record, and a ledger can record a saving the workforce does not feel. Both are real findings and both need reporting.
Four questions to answer before scaling either use case
| Question before scale | Ambient documentation | Inbox drafting |
|---|---|---|
| Should the underlying output exist in its current form? | Can the note be shorter and clinically sufficient? | Can routing, scope, and demand be redesigned before drafting? |
| What new work is created? | Consent, review, correction, attribution, note retrieval | Reading drafts, personalization, safety review, escalation |
| What is the patient outcome? | Attention, trust, documentation quality, diagnostic reflection | Response quality, access, resolution, unnecessary visits |
| What is the stop rule? | Unsafe omission, unacceptable edit burden, note bloat, specialty mismatch | Hallucinated advice, no net workload benefit, inequitable access |
Redesign opportunity
The highest-value documentation strategy may combine requirement elimination, shorter note standards, team-based work, structured data capture, and selective AI support.
Automating every current documentation expectation risks making note bloat cheaper rather than making documentation better.
Human accountability cannot be automated
Human-in-the-loop is a location in a diagram, not a reliable control
Healthcare AI governance often relies on a reassuring phrase: a clinician remains in the loop. That phrase does not specify whether the clinician has time, expertise, independent information, authority, interface support, or a realistic opportunity to disagree. Oversight imposed as a final click can become ceremonial verification, especially when automation volume is high and errors are uncommon enough to weaken vigilance.
Gaube and colleagues studied 138 radiologists and 127 internal or emergency medicine physicians who reviewed chest radiographs with accurate or inaccurate advice, labeled as coming from either AI or a radiologist. Diagnostic performance was shaped by advice accuracy rather than by the purported source of the advice. Among internal or emergency medicine physicians, 41.73% followed all incorrect advice they received; 27.54% of radiologists did the same.
Physicians who followed every piece of incorrect advice
Share of each group, with the underlying participant count printed on each bar.
Chart source. Gaube et al. (2021) as reported in the research report. Percentages are published. The participant counts are arithmetic on the published group sizes and are labeled as derived: 53 of 127 reproduces 41.73% and 38 of 138 reproduces 27.54%, so both published rates resolve to whole participants.
of all 265 participants followed every incorrect recommendation. Derived here from the two published rates.
What the design can and cannot support
The study was experimental and involved eight simulated radiograph cases. It cannot estimate real-world error rates. It does demonstrate why accountability cannot be reduced to nominal human review, which is the only inference the report draws from it.
The finding that unsettles the standard control
Performance tracked the accuracy of the advice, not its labeled source. Telling clinicians that a recommendation came from a colleague rather than a model did not protect them.
Eight dimensions that make oversight real
The report replaces the phrase with a specification. Each dimension below is a question that has an answer or does not; there is no partial credit in the source table.
| Accountability dimension | Required specification |
|---|---|
| Decision right | Who can accept, reject, modify, or escalate AI output? |
| Duty to review | Which elements must be independently verified, and under what time standard? |
| Competence | What task expertise and AI-specific training are required? |
| Capacity | Is protected time available for review and exception management? |
| Traceability | Can the organization reconstruct data, output, human action, rationale, and outcome? |
| Escalation | Who receives disagreement, uncertainty, drift, safety events, and patient concerns? |
| Authority to stop | Who can pause the workflow without vendor or executive delay? |
| Liability and disclosure | How are responsibility, patient notice, consent, and contestability addressed? |
The six-condition oversight test
The report states that meaningful human oversight requires independence, information, competence, time, authority, and feedback, and that if one is absent the organization should treat the control as weak and add system-level safeguards. That is a rule with no partial credit, so this tool does not award any.
Oversight strength check
Six conditions. The report’s own rule is applied literally: one absent condition makes the control weak, regardless of how many others are present.
Oversight conditions present
0 of 6
Mark each condition that is genuinely in place.
No score is reported when a condition is missing, because a control that can be outweighed by other strengths is not a control. This mirrors how the AWRI treats its stop gates.
Control design principle
Meaningful human oversight requires independence, information, competence, time, authority, and feedback.
If one is absent, the organization should treat the control as weak and add system-level safeguards.
AI can industrialize inequity
Bias is often a workflow and objective problem before it is a model problem
AI learns from recorded care, but recorded care is not the same as need. It reflects who gained access, which services were paid for, what was documented, which populations were studied, and how organizations defined success. A model can therefore be accurate for a target that reproduces unequal systems.
Obermeyer and colleagues examined a widely used population health algorithm that used cost as a proxy for health need. Because less money was spent on Black patients with the same level of illness, the algorithm assigned them a lower risk. Correcting the objective would have increased the share of Black patients receiving additional help from 17.7% to 46.5%. The defect was not simply statistical. It was a mis-specified operational objective embedded in a resource-allocation workflow.
What changing the objective, and nothing else, would have done
Share of patients receiving additional help under the cost proxy and under a corrected objective. The model architecture is unchanged between the two bars.
Chart source. Obermeyer et al. (2019) as reported in the research report. Both values are published. The 28.8 percentage-point difference and the 2.63-fold ratio are arithmetic and are labeled as derived.
increase in the share receiving additional help, a gain of 28.8 percentage points, from a change to the objective rather than the algorithm.
Why this belongs in a workflow report
The algorithm was technically accurate. It predicted cost well. The failure was in choosing cost as the thing worth predicting, and that choice was made in the design of a resource-allocation workflow, not in a training run.
Stated limitation
The report records the study’s own boundary: one class of algorithm and one operational context.
Six sources of inequity, each with a workflow manifestation
The report’s structure is deliberate. Every row names where the problem becomes visible in the work, not only where it originates in the data.
| Source of inequity | Workflow manifestation | Governance response |
|---|---|---|
| Access bias | Training data contains care received, not care needed | Measure unmet need and eligibility outside historical utilization |
| Proxy bias | Cost, attendance, or documentation substitutes for clinical need | Validate target meaning with clinicians, patients, and equity leaders |
| Representation bias | Subgroups are missing or too small | Local subgroup validation, uncertainty reporting, deployment limits |
| Automation bias | Users follow confident recommendations | Independent review, uncertainty display, counterfactual testing, training |
| Capacity bias | Only well-resourced patients can act on recommendations | Measure completion, burden, language, transport, digital, and financial barriers |
| Feedback-loop bias | AI-directed care changes future data and reinforces allocation | Monitor the longitudinal distribution of benefit, harm, and opportunity |
Equity gate
A deployment should not pass because subgroup accuracy is acceptable.
Leaders must test whether the target, workflow, available resources, and downstream actions distribute opportunity and burden fairly.
Objective validity check
The report’s central equity argument is about the target variable, so this check asks about the target rather than about model performance. It returns a reading, not a score.
What is your model actually predicting?
Answer for the workflow you are considering. The check flags the specific failure mode each answer maps to in the table above.
Equity risks flagged
0 of 6
Each flagged answer names a source of inequity from the table.
The report states that a deployment should not proceed when local equity evaluation is infeasible for the intended population. That condition is reported here as a stop rather than as a deduction, consistent with how it appears in the AWRI stop gates.
Proposed AI Workflow Readiness Index
A 100-point internal governance screen with noncompensable safety gates
The AWRI is an original, provisional instrument for comparing candidate workflows and documenting governance readiness.
It is not a clinical score, a procurement ranking, or a validated prediction model. Organizations should test reliability, validity, gaming risk, and decision usefulness before using it for incentives or external comparison. The calculator below reproduces the published weights exactly and enforces the published stop gates as stops, not as deductions.

AWRI calculator
Score each domain out of its published weight. The readout reports the band, and separately reports whether the workflow may proceed, because those are two different questions in the source model.
Should this work exist, and is AI preferable to a simpler redesign?
Is the target process mapped, measured, and reliable?
Are live local data suitable for the intended target?
Are decision rights, review, escalation, and stop authority explicit?
Can intended users safely integrate the output under a realistic workload?
Can the organization detect loss of value and act?
The two stop gates that are not scores
Either condition stops production regardless of the total, in the report’s own wording.
Points scored against available weight in each domain. The two gated domains are drawn with their thresholds marked.
AWRI composite
Recommended action at this band
Do not automate. Eliminate and redesign the process.
The published weights and bands
| Domain | Weight | Core test |
|---|---|---|
| Necessity | 15 | Should this work exist, and is AI preferable to a simpler redesign? |
| Workflow stability | 20 | Is the target process mapped, measured, and reliable? |
| Data fitness | 15 | Are live local data suitable for the intended target? |
| Accountability | 20 | Are decision rights, review, escalation, and stop authority explicit? |
| Human factors | 15 | Can intended users safely integrate the output under a realistic workload? |
| Lifecycle monitoring | 15 | Can the organization detect loss of value and act? |
| AWRI range | Interpretation | Recommended action |
|---|---|---|
| 0 to 39 | Not ready | Do not automate. Eliminate and redesign the process. |
| 40 to 59 | Conditional | Address specific readiness gaps; no production use. |
| 60 to 79 | Pilot-ready | Silent mode or bounded pilot with active oversight. |
| 80 to 100 | Scale consideration | Scale only if outcome and guardrail evidence remain favorable. |
Stop gates
Regardless of total score, do not proceed to production if accountability is below 12 of 20, lifecycle monitoring is below 9 of 15, a high-severity hazard lacks control, or local equity evaluation is infeasible for the intended population.
Derived finding
The band table and the stop gates are two different instruments
The report states both rules plainly, and they are not in conflict when read carefully: the band is an interpretation of readiness, and the gates are a permission to proceed. In practice, a governance committee reading a single number can easily treat the top band as approval. The enumeration below shows how far apart the two rules can sit.
ways to score the six domains in whole points, from 16 x 21 x 16 x 21 x 16 x 16.
score 80 or above, which the published table labels Scale consideration.
of those, 17.04%, fail the accountability gate, the lifecycle gate, or both.
score 80 or above while failing both scored gates at once.
Highest total reachable while failing a gate
Each bar is the maximum composite a workflow can reach while the named gate is failing. The dashed line is the lower edge of the Scale consideration band.
Author-derived. Computed by exhaustive enumeration of every whole-point combination of the six published domain weights against the two published scored stop gates. Verified independently before publication. The report states both rules; the enumeration is arithmetic on them, not a new claim.
The headline number
A workflow can score 91 of 100 with accountability at 11 of 20, and 93 of 100 with lifecycle monitoring at 8 of 15. A workflow failing both gates at once can still reach 84, which sits four points inside the top band.
The reverse also holds
A workflow scoring only 21 passes both scored gates, with accountability at 12 and lifecycle at 9 and every other domain at zero. Passing the gates is not readiness, and reaching the band is not permission. The calculator above therefore reports the band and the verdict as separate lines and never blends them into one number.
Design consequence adopted here: when a gate fails, the calculator reports no recommendation to proceed at all rather than a reduced one. A gate that can be outweighed by strength elsewhere is not a gate.
Proposed Industrialized Dysfunction Risk model
A conceptual screening equation, not a validated prediction instrument
The model multiplies four domains: workflow waste, automation scale, propagation speed, and opacity. Its published decision rule is short and unambiguous: if two domains are high, redesign before scale. The report is explicit that the multiplicative form is illustrative and that this is not a calibrated risk score.

Dysfunction risk screener
Score each domain from 1 to 5. The screener reports the published rule and the illustrative product separately, because they do not always agree, and it says so when they do not.
Unnecessary steps, rework, weak controls.
Patients, encounters, sites, decisions touched.
How quickly output triggers downstream action.
How hard failure is to detect and contest.
Domain scores against the high threshold of 4 that the published rule uses. Bars at or above the marked line count toward the rule.
Published decision rule
Proceed
Fewer than two domains are high
Which reading governs
The rule governs. The product is shown because the report draws the model as a multiplication, but the report states the decision in terms of the rule, and only the rule is offered as a screening decision.
The product is reported without units and never converted into a percentage, a band, or a risk probability, because the report gives it no scale.
Derived finding
The multiplication and the rule rank differently
Scored 1 to 5, the four domains produce 625 profiles. If the product were used as a ranking, it would order those profiles differently from the published rule. The two disagree on 174 of the 625, which is 27.84%.
whole-point combinations of the four domains.
have two or more domains at 4 or above, 52.48% of all profiles.
profiles with a product at or above the 75th percentile of 100.
of profiles, 174 of 625, are flagged by one reading and not the other.
Two profiles the two readings rank in opposite order
Same four domains, same model. As a ranking the product puts the left profile above the right; the published rule stops only the right.
Author-derived. Computed by exhaustive enumeration of all 625 whole-point profiles against the published two-high rule and the multiplicative product. Verified independently before publication.
The clearest case
A profile of 3-3-3-3 produces a product of 81 and is never flagged, because no domain reaches 4. A profile of 5-5-1-1 produces a product of 25, less than a third as large, and is always flagged. On the product ranking the first looks worse; on the published rule only the second is stopped.
Why the rule is the better instrument here
A multiplication rewards a low score in any single domain very heavily, so a workflow with one contained domain can mask two severe ones. The two-high rule is insensitive to that masking, which is exactly the failure mode the report is warning about. The screener therefore treats the product as illustration and the rule as the decision.
Range check: flagged profiles span products from 16 to 625, unflagged profiles span 1 to 135, and 179 unflagged profiles carry a product above the lowest flagged one. The two ranges overlap heavily, so the product cannot be used to reconstruct the rule.
Stated limitation
The Industrialized Dysfunction Risk model is conceptual, and its multiplicative form is illustrative.
Nothing in this section should be read as validation of the model. The enumeration is arithmetic on the report’s own published rule and figure, offered so that a screening committee does not treat the drawing of a multiplication as a ranking instrument.
Governance for the human-plus-AI system
The board governs the clinical operating model, while management runs a continuous control loop

The 2025 JAMA Summit report emphasized holistic, continuous, multistakeholder lifecycle management and noted that widely adopted tools often have limited evaluation of health effects. The report also observed that effectiveness depends on the human-computer interface, user training, and setting, and that no overarching accountability structure spans the many stakeholders. FUTURE-AI similarly organizes trustworthy AI across fairness, universality, traceability, usability, robustness, and explainability, with practices spanning design through monitoring.
Principles are not an operational path
A 2026 scoping review of 77 healthcare AI governance frameworks found that only 10, or 13.0%, included guiding principles, assessment methods, lifecycle stages, and oversight mechanisms together. Oversight mechanisms appeared in only 15 frameworks. This gap explains why organizations can possess extensive principles and still lack an operational path to stop a harmful tool.
Coverage across 77 governance frameworks
Frameworks containing all four elements, and frameworks containing an oversight mechanism at all.
Chart source. Wang et al. (2026) as reported in the research report. Counts and the 13% figure are published; the 19.5% oversight share is arithmetic on 15 of 77 and is labeled as derived.
13.0% covered principles, assessment methods, lifecycle stages, and oversight mechanisms together.
19.5%. Four in five frameworks describe governance without describing who acts.
Stated limitation
The report records the review’s own boundary: framework presence does not equal real-world adoption. A framework counted here may never have governed a deployment.
Six governance layers, and the question each one owns
Select a layer to see the accountability it holds and the questions it must be able to answer.
Primary accountability
Board
Risk appetite, strategy, patient, and workforce outcomes
Questions that must be answered
What this layer must be able to say
Which uses are prohibited, material, or board-reportable? What value and harm thresholds apply?
| Governance layer | Primary accountability | Questions that must be answered |
|---|---|---|
| Board | Risk appetite, strategy, patient, and workforce outcomes | Which uses are prohibited, material, or board-reportable? What value and harm thresholds apply? |
| Executive AI council | Portfolio, investment, enterprise standards, stop authority | Which workflows pass readiness gates? Who owns cross-functional tradeoffs? |
| Clinical and operational owners | End-to-end pathway performance | Does the combined process improve patient outcomes and total work? |
| Model and data owners | Data lineage, model performance, drift, technical change | Are data and model behavior within validated bounds? |
| Safety, quality, privacy, equity | Independent assurance and incident review | Are harms detectable, contestable, and correctable? |
| Frontline users and patients | Use, feedback, refusal, escalation, lived experience | Can people understand, challenge, and report problems? |
Nondelegable responsibility
A vendor may maintain the software, and a clinician may review the output, but the health system remains responsible for the care pathway it chooses to operate.
Accountability cannot be outsourced through a contract or transferred through a click.
Stakeholder benefit, burden, and the protection each one requires
| Stakeholder | Potential benefit | Potential industrialized burden | Required protection |
|---|---|---|---|
| Patients | Faster access, clearer communication, more clinician attention | Opaque decisions, digital burden, longer records, inequitable service | Notice, consent where appropriate, contestability, accessible alternatives |
| Clinicians | Less drafting, better prioritization, decision support | Verification, alert fatigue, liability ambiguity, deskilling | Protected review time, training, authority, outcome feedback |
| Nurses and staff | Reduced routing and repetitive entry | Exception queues, hidden triage, correction work | Workload measurement, role clarity, escalation, staffing |
| Operational leaders | Visibility and capacity improvement | False certainty, vendor dependence, benefit overstatement | Independent metrics, audit rights, retirement criteria |
| Communities | Expanded reach and earlier intervention | Historical bias scaled across populations | Community input, equity measurement, alternative access paths |
Workforce compact
Before deployment, leaders should state how time savings will be used, how new verification and exception work will be staffed, and how workers can report unsafe or burdensome effects without penalty.
AI may reduce some forms of cognitive and clerical burden, but savings are not automatically returned to clinicians or patients. Organizations may convert faster documentation into higher visit volume, shift verification to lower-paid staff, or transfer self-service work to patients.
The 2026 American Hospital Association workforce scan states that AI works best when paired with redesigned processes. The report treats this as industry context rather than peer-reviewed evidence, and notes that it aligns with the peer-reviewed implementation literature.
The board dashboard
Measure value, work, and harm
Adoption is not the outcome. Activity counts can reward the industrialization of the wrong process.
Measures should be stratified by site, service, population, urgency, role, and time when sample size permits. A favorable average can conceal failure at a single hospital, in a burdensome specialty, or within an underserved population. Trend limits and decision thresholds should be established before leaders know whether results are favorable.
Ten domains a board should see
| Domain | Core measure | Definition and signal | Example board trigger |
|---|---|---|---|
| Patient value | Outcome delta | Risk-adjusted change in the patient outcome that the workflow exists to improve | No improvement after the agreed learning period |
| Total work | Net work per completed episode | All human minutes across creation, review, correction, exception, and downstream action | Task time falls, but total work rises |
| Action yield | Useful actions per 100 AI outputs | Outputs that produce a necessary, timely, completed action | Yield below threshold or declining |
| Failure demand | Rework and exceptions per 100 episodes | Work was created because the process failed or the output was unusable | High-severity exception without owner |
| Human reliability | Override, disagreement, and review fidelity | Patterns of acceptance, correction, nonuse, and missed review | Blind acceptance or ceremonial review |
| Equity | Benefit and burden gaps | Outcome, access, false result, workload, and completion differences | Material unexplained subgroup gap |
| Safety | AI-related harm and near-miss rate | Severity-weighted events with causal contribution review | Sentinel event or repeated high-risk precursor |
| Drift | Performance and context stability | Model, data, workflow, staffing, and clinical-practice change | Validated bound exceeded |
| Resilience | Downtime and recovery | Ability to operate safely when AI is unavailable or degraded | Unsafe fallback or prolonged recovery |
| Lifecycle control | Open corrective actions | Overdue safety, equity, vendor, model, or workflow actions | Material action overdue |
Vanity metric translator
The report pairs five commonly reported metrics with what should replace them. Select the metric your reporting pack currently leads with.
What your current metric cannot tell the board
Each replacement answers a question the original cannot, and each names the failure mode the original permits.
What it permits
Number of AI tools deployed
Replace with
Number of workflows with demonstrated patient value and controlled risk
| Avoid | Replace with |
|---|---|
| Number of AI tools deployed | Number of workflows with demonstrated patient value and controlled risk |
| Monthly active users | Appropriate use, nonuse, override, and outcome by context |
| Hours saved estimated by vendor | Observed network removed and capacity actually released |
| Model accuracy alone | Human-plus-AI pathway performance and patient outcome |
| No reported incidents | Detection sensitivity, near misses, speaking-up climate, and corrective action closure |
Metric dictionary
Proposed operational definitions for local governance and benefits realization. Filter by the direction each metric should move, or by owner.
| Metric | Proposed operational definition | Direction | Owner | Guardrail |
|---|---|---|---|---|
| Low-value work eliminated | Count and minutes of steps retired before automation, verified in target state | Higher | Process owner | No loss of required clinical or regulatory control |
| Net work per completed episode | Total human minutes across all roles from demand to resolved outcome | Lower | Operations | Patient outcome and experience were maintained or improved |
| Action yield | AI outputs producing a necessary, timely, completed action divided by total outputs | Higher | Clinical owner | Monitor missed need and overtreatment |
| Failure demand | Work generated because output, routing, capacity, or process failed | Lower | Operations and quality | Do not suppress reporting |
| Review fidelity | Required review elements completed with evidence of independent judgment | Higher | Clinical owner | Avoid click-based proxy alone |
| AI disagreement rate | Outputs materially corrected or rejected by authorized users | Contextual | Model owner | Investigate both abrupt change and implausibly low rates |
| Exception aging | Open exception cases by time, severity, and ownership state | Lower | Operational owner | Separate patient choice and clinically appropriate delay |
| Equity gap | Absolute and relative difference in benefit, burden, or harm across groups | Lower | Quality and equity | Report data completeness and small-cell instability |
| Drift response time | Time from threshold breach to triage, containment, decision, and closure | Lower | Technical and clinical owners | Severity-based target |
| Released capacity realized | Observed clinical or access capacity created and intentionally redeployed | Higher | Executive sponsor | Do not count estimated vendor hours |
| AI-attributable incident rate | Events in which AI output, interface, data, or workflow contributed | Lower | Safety | Include near misses and sociotechnical causes |
| Recovery reliability | Safe performance during downtime and time to restore validated workflow | Higher | IT and operations | Test fallback under realistic load |
Minimum efficiency claim
Do not claim efficiency from reduced task time alone.
Demonstrate net change in total work per completed episode, including verification, exceptions, rework, maintenance, and downstream utilization, then show how released capacity was actually used.
Implementation roadmap
A staged operating model that begins with elimination and evidence, not enterprise licensing
The first portfolio should favor workflows with a clear patient outcome, high avoidable burden, measurable baseline, stable data, bounded action, strong ownership, and reversible deployment. Avoid beginning with high-autonomy, high-severity decisions, workflows with weak data provenance, or processes whose demand and exception patterns are not understood.
Four horizons, and the decision each one authorizes
Select a horizon to see the leadership actions, required deliverables, and the scale decision the report permits at that point.
Leadership actions
0 to 90 days
Required deliverables
What must exist by the end
Scale decision permitted
What you may authorize
| Time horizon | Leadership actions | Required deliverables | Scale decision |
|---|---|---|---|
| 0 to 90 days | Create portfolio governance; identify high-burden workflows; freeze scale decisions lacking ownership; map current state | AI use inventory; prohibited uses; process maps; baseline outcomes; initial AWRI | Select no more than a small number of workflows for redesign |
| 3 to 6 months | Eliminate low-value work; design target state; specify roles, measures, hazards, and vendor requirements | Target-state map; failure-mode analysis; RACI; data and equity plan; test protocol | Authorize silent mode only after the gates pass |
| 6 to 12 months | Run silent trials and bounded pilots; observe real work; train users; test downtime and escalation | Local validation; usability results; incident log; total-work study; patient feedback | Proceed only with a favorable outcome and guardrail evidence |
| 12 to 24 months | Scale by site with controlled learning; monitor drift; renegotiate or retire low-value tools | Site readiness; lifecycle dashboard; audit trail; benefits realization; retirement plan | Expansion remains reversible and evidence-dependent |
Discussion
The decisive question is not whether healthcare should use AI, but what operating system AI should amplify.
The central hypothesis is supported by a consistent pattern: healthcare AI value emerges from the interaction of model, workflow, people, resources, interface, accountability, and learning. This pattern appears in foundational systems theory, contemporary implementation reviews, human factors experiments, generative AI field studies, governance consensus, and contrasting real-world outcomes. The evidence does not justify rejecting AI. It justifies rejecting automation as a substitute for improvement.
The title’s force comes from a genuine asymmetry. Once automation is embedded, scale and sunk cost create pressure to preserve it. Users adapt, policies are written around it, vendors integrate it, and metrics begin to reward activity. A flawed manual workflow is visible in queues and complaints. A flawed automated workflow can appear efficient because the activity is rapid and its costs are dispersed across review, exceptions, downstream care, patient effort, and rare harms.
The guard against overcorrection
Workflow-first sequencing should not become perfectionism. Processes in complex care will never be fully stable, and waiting for complete certainty would block useful innovation. The disciplined alternative is bounded uncertainty: define the problem, remove obvious waste, design responsibility, test on live local data without influence, deploy reversibly, measure total effects, and retain authority to stop.
Conclusion
Do not automate accumulated dysfunction and call it transformation.
Artificial intelligence can improve healthcare. The evidence supports real gains in documentation burden, prediction-enabled response, and selected clinical processes. But those gains are conditional. AI does not decide whether a task is necessary, whether a proxy is morally and clinically valid, whether an exception has an owner, whether a reviewer has time, or whether released capacity returns to patients. Leaders must make those decisions.
The workflow-first sequence is straightforward: eliminate low-value work, redesign the complete pathway, assign clinical and operational accountability, and automate only the validated process. This sequence changes the leadership conversation from how quickly a tool can be deployed to whether the organization is prepared to amplify the process it has built.
Final charge to leaders
Before asking what AI can do, ask what work should disappear, what care process should exist, who owns the result, and what evidence would justify scale.
Transformation begins before automation. The central warning is therefore also an opportunity. AI will industrialize the operating system it enters. Healthcare leaders can allow accumulated dysfunction to scale, or they can use the moment to remove work, rebuild accountability, and design safer care. The technology will amplify either choice.
Claims ledger
What the evidence supports, and what it does not
The report grades its own claims rather than presenting them as uniform findings. Two of the six are explicitly not established, including a weaker form of the report’s own title.
| Claim | Evidence position | Leadership interpretation |
|---|---|---|
| AI will industrialize every broken workflow | Not established and too universal | Use as a warning about scale, not a deterministic forecast |
| Workflow design determines realized AI value | Strong convergent support | Treat workflow as part of the intervention and investment case |
| Elimination should precede automation | Normatively strong; direct comparative trials limited | Require a formal necessity and de-implementation gate |
| Human review makes AI safe | Not supported as a standalone control | Design meaningful oversight and system safeguards |
| Task time saved equals efficiency | Contradicted by review and field evidence | Measure total work and capacity released |
| Continuous monitoring is required | Strong consensus and implementation support | Fund lifecycle operations as recurring clinical infrastructure |
Strategic conclusion
The most advanced health system will not be the one with the most AI.
It will be the one that can reliably decide which work to eliminate, which process to redesign, which decision to augment, which risk to accept, and which tool to stop.
Evidence matrix
The twelve studies most directly informing the hypothesis, each with the limitation the report records against it. Filter by design.
| Study | Design and setting | Finding used | Principal limitation |
|---|---|---|---|
| Wenderott et al., 2024 | Systematic review and meta-analysis of 48 real-world imaging studies | Task-time reductions are often reported; pooled effects are nonsignificant; workload translation is rarely examined | Heterogeneous tasks, tools, and study quality |
| Khan et al., 2024 | Systematic review of 25 AI implementation frameworks | Plan emphasized more than Do, Study, and especially Act | Framework content does not establish implementation effectiveness |
| Preti et al., 2024 | Systematic review of 34 empirical ML implementations | Workflow, design, infrastructure, engagement, leadership, and evaluation recur | Published implementations may overrepresent success |
| Rahimi et al., 2024 | Systematic review of 26 hospital AI implementation studies | 28 enablers and 18 barriers; 81% of included papers rated poor quality | Limited high-quality real-world evidence |
| Nair et al., 2024 | 38 cases plus 69 stakeholder interviews | Twelve implementation concepts span planning, use, and sustainment | Mixed evidence sources and qualitative interpretation |
| Wong et al., 2021 | External validation in 38,455 hospitalizations | AUC 0.63; 18% alerted; 67% of sepsis missed | Retrospective single-system validation; hypothetical alert burden |
| Boussina et al., 2024 | Before-and-after quasi-experimental deployment in two EDs | Associated with lower mortality and higher bundle compliance within an integrated workflow | Nonrandomized; benefit varied by hospital |
| Duggan et al., 2025 | Single-group pre-post study of 46 clinicians | Less note and after-hours time, longer notes, mixed feedback | Short, opt-in, single-system study |
| Tai-Seale et al., 2024 | Randomized waiting-list QI study with contemporary controls | More read time, no reply-time reduction, longer replies | One early tool in an academic system |
| Gaube et al., 2021 | Web experiment with 265 physicians | Many physicians followed inaccurate advice regardless of the source | Eight simulated radiograph cases |
| Obermeyer et al., 2019 | Analysis of population health algorithm and clinical data | Cost proxy encoded racial inequity in access to extra care | One class of algorithm and operational context |
| Wang et al., 2026 | Scoping review of 77 AI governance frameworks | Only 13% covered principles, assessment, lifecycle, and oversight | Framework presence does not equal real-world adoption |
Limitations and research agenda
The practical framework precedes direct comparative evidence and should be tested accordingly.
The report’s hypothesis is not supported by randomized comparisons of E-R-A-A against product-first implementation. Many cited studies are observational, pre-post, quality improvement, simulation, or review designs. Real-world AI interventions are bundles, making it difficult to isolate the model from training, interface, staffing, communication, and concurrent improvement. Publication bias may favor successful implementations. Vendor evolution also makes results time-dependent.
The original models are provisional. E-R-A-A has face validity from implementation and systems evidence but requires empirical evaluation. The Industrialized Dysfunction Risk model is conceptual, and its multiplicative form is illustrative. The AWRI weights and thresholds are proposed for internal testing and should not be treated as psychometrically established.
| Research priority | Suggested design | Primary outcomes |
|---|---|---|
| Workflow-first versus product-first sequencing | Cluster-randomized or stepped-wedge implementation trial | Patient outcome, total work, time to value, safety, sustainment |
| De-implementation before generative AI | Factorial trial of requirement reduction and AI support | Documentation burden, note quality, note length, diagnostic reflection |
| Meaningful human oversight | Simulation plus prospective field evaluation | Error detection, automation bias, review time, escalation fidelity |
| Network migration | Time-motion, EHR logs, task mining, and qualitative observation | Work created, shifted, eliminated, and released by role |
| AWRI validation | Multisite prospective cohort of AI projects | Reliability, predictive validity, responsiveness, gaming, decision utility |
| Equity in workflow outcomes | Subgroup and distributional impact evaluation | Access, false results, completion, burden, benefit, contestability |
| Retirement science | Comparative case study of pause and decommission decisions | Harms avoided, switching cost, organizational learning |
Publication pathway
For journal submission, select an article type and reporting standard, conduct a registered systematic or realist review if required, complete dual screening and risk-of-bias appraisal, and empirically test E-R-A-A or the AWRI in multiple health systems.
Verification ledger for this dashboard
Every figure this dashboard derives, and whether it reproduces the report’s published value.
| Check | Result | Note |
|---|---|---|
| PDSA percentages against a denominator of 25 | Reproduce exactly | 84%, 60%, 52%, and 24% resolve to 21, 15, 13, and 6 frameworks, all whole numbers, confirming 25 as the denominator throughout. |
| Automation bias percentages against group sizes | Reproduce exactly | 53 of 127 gives 41.73% and 38 of 138 gives 27.54%. Both published rates resolve to whole participants. |
| Share of task-time studies reporting a reduction | Resolves to 22 of 33 | 22 of 33 is 66.67% and rounds to 67%. 23 of 33 is 69.7% and does not. The count is therefore determined. |
| COMPOSER absolute and relative results | Internally consistent | The paired figures imply a baseline bundle compliance of exactly 50.0% and a baseline mortality near 11.2%. |
| Implied sepsis prevalence in the external validation | Determined only to a range | Because the three published rates are rounded to whole percentage points, the implied prevalence is 6.01% to 7.12%, centred near 6.55%. Reported as a range rather than a point. |
| AWRI band table against its own stop gates | Two separate instruments | 39,088 of the 229,390 top-band score vectors fail a scored gate. The dashboard therefore reports band and verdict on separate lines. |
| Dysfunction model product against its two-high rule | Disagree on 27.84% | 174 of 625 profiles are flagged by one reading and not the other, so the product is presented as illustration only. |
References
Peer-reviewed and authoritative sources cited in the report. DOI and source links are live.
- Adams, R., et al. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28, 1455-1460. doi.org/10.1038/s41591-022-01894-0
- Angus, D. C., Khera, R., Lieu, T., et al. (2025). AI, health, and health care today and tomorrow: The JAMA Summit report on artificial intelligence. JAMA, 334(18), 1650-1664. doi.org/10.1001/jama.2025.18490
- Arndt, B. G., Beasley, J. W., Watkinson, M. D., et al. (2017). Tethered to the EHR: Primary care physician workload assessment using EHR event log data and time-motion observations. Annals of Family Medicine, 15(5), 419-426. doi.org/10.1370/afm.2121
- Beede, E., Baylor, E., Hersch, F., et al. (2020). A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 1-12. doi.org/10.1145/3313831.3376718
- Boussina, A., Shashikumar, S. P., Malhotra, A., et al. (2024). Impact of a deep learning sepsis prediction model on quality of care and survival. npj Digital Medicine, 7, 14. doi.org/10.1038/s41746-023-00986-6
- Cabitza, F., Rasoini, R., & Gensini, G. F. (2017). Unintended consequences of machine learning in medicine. JAMA, 318(6), 517-518. doi.org/10.1001/jama.2017.7797
- Carayon, P., Schoofs Hundt, A., Karsh, B. T., et al. (2006). Work system design for patient safety: The SEIPS model. Quality & Safety in Health Care, 15(Suppl 1), i50-i58. doi.org/10.1136/qshc.2005.015842
- Duggan, M. J., Gervase, J., Schoenbaum, A., et al. (2025). Clinician experiences with ambient scribe technology to assist with documentation burden and efficiency. JAMA Network Open, 8(2), e2460637. doi.org/10.1001/jamanetworkopen.2024.60637
- Gaube, S., Suresh, H., Raue, M., et al. (2021). Do as AI says: Susceptibility in deployment of clinical decision aids. npj Digital Medicine, 4, 31. doi.org/10.1038/s41746-021-00385-9
- Garcia, P., et al. (2024). Artificial intelligence-generated draft replies to patient inbox messages. JAMA Network Open, 7(3), e243201. doi.org/10.1001/jamanetworkopen.2024.3201
- Greenhalgh, T., Wherton, J., Papoutsi, C., et al. (2017). Beyond adoption: A new framework for theorizing and evaluating nonadoption, abandonment, and challenges to scale-up, spread, and sustainability of health and care technologies. Journal of Medical Internet Research, 19(11), e367. doi.org/10.2196/jmir.8775
- He, J., Baxter, S. L., Xu, J., Xu, J., Zhou, X., & Zhang, K. (2019). The practical implementation of artificial intelligence technologies in medicine. Nature Medicine, 25, 30-36. doi.org/10.1038/s41591-018-0307-0
- Holmgren, A. J., Hendrix, N., Maisel, N., Everson, J., & Bazemore, A. (2024). Electronic health record usability, satisfaction, and burnout for family physicians. JAMA Network Open, 7(8), e2426956. doi.org/10.1001/jamanetworkopen.2024.26956
- Kamel Rahimi, A., Pienaar, O., Ghadimi, M., et al. (2024). Implementing AI in hospitals to achieve a learning health system: Systematic review of current enablers and barriers. Journal of Medical Internet Research, 26, e49655. doi.org/10.2196/49655
- Khan, A. A., et al. (2024). Implementation frameworks for artificial intelligence translation into health care practice: A systematic review. PLOS Digital Health, 3, e0000514. doi.org/10.1371/journal.pdig.0000514
- Kwong, J. C. C., Nickel, G. C., Wang, S. C. Y., & Kvedar, J. C. (2024). Integrating artificial intelligence into healthcare systems: More than just the algorithm. npj Digital Medicine, 7, 52. doi.org/10.1038/s41746-024-01066-z
- Lekadir, K., Frangi, A. F., Porras, A. R., et al. (2025). FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ, 388, e081554. doi.org/10.1136/bmj-2024-081554
- Liu, T. L., Hetherington, T. C., Dharod, A., et al. (2024). Does AI-powered clinical documentation enhance clinician efficiency? A longitudinal study. NEJM AI, 1(12), AIoa2400659. doi.org/10.1056/AIoa2400659
- Nair, M., Svedberg, P., Larsson, I., & Nygren, J. M. (2024). A comprehensive overview of barriers and strategies for AI implementation in healthcare: Mixed-methods design. PLOS ONE, 19(8), e0305949. doi.org/10.1371/journal.pone.0305949
- Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453. doi.org/10.1126/science.aax2342
- Preti, L. M., Ardito, V., Compagni, A., Petracca, F., & Cappellaro, G. (2024). Implementation of machine learning applications in health care organizations: Systematic review of empirical studies. Journal of Medical Internet Research, 26, e55897. doi.org/10.2196/55897
- Sendak, M. P., Ratliff, W., Sarro, D., et al. (2020). Real-world integration of a sepsis deep learning technology into routine clinical care: Implementation study. JMIR Medical Informatics, 8(7), e15182. doi.org/10.2196/15182
- Sinsky, C., Colligan, L., Li, L., et al. (2016). Allocation of physician time in ambulatory practice: A time and motion study in 4 specialties. Annals of Internal Medicine, 165(11), 753-760. doi.org/10.7326/M16-0961
- Sittig, D. F., & Singh, H. (2010). A new sociotechnical model for studying health information technology in complex adaptive healthcare systems. Quality & Safety in Health Care, 19(Suppl 3), i68-i74. doi.org/10.1136/qshc.2010.042085
- Sriharan, A., Sekercioglu, N., Mitchell, C., et al. (2024). Leadership for AI transformation in health care organizations: Scoping review. Journal of Medical Internet Research, 26, e54556. doi.org/10.2196/54556
- Tai-Seale, M., Baxter, S. L., Vaida, F., et al. (2024). AI-generated draft replies integrated into health records and physicians’ electronic communication. JAMA Network Open, 7(4), e246565. doi.org/10.1001/jamanetworkopen.2024.6565
- van de Sande, D., Chung, E. F. F., Oosterhoff, J., et al. (2024). To warrant clinical adoption, AI models require a multi-faceted implementation evaluation. npj Digital Medicine, 7, 51. doi.org/10.1038/s41746-024-01064-1
- Vasey, B., Nagendran, M., Campbell, B., et al. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine, 28, 924-933. doi.org/10.1038/s41591-022-01772-9
- Wang, A., Freeman, S., & Magrabi, F. (2026). Governance for safe and responsible AI in healthcare organizations: A scoping review of frameworks. npj Digital Medicine, 9. doi.org/10.1038/s41746-026-02679-2
- Wenderott, K., Krups, J., Weigl, M., & Wooldridge, A. (2024). Effects of artificial intelligence implementation on efficiency in medical imaging: A systematic literature review and meta-analysis. npj Digital Medicine, 7, 265. doi.org/10.1038/s41746-024-01248-9
- Wenderott, K., Krups, J., Weigl, M., & Wooldridge, A. (2026). Navigating the complexity of artificial intelligence implementation in healthcare: Ten recommendations grounded in implementation science. BMJ Open Quality, 15, e003639. doi.org/10.1136/bmjoq-2025-003639
- Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine, 181(8), 1065-1070. doi.org/10.1001/jamainternmed.2021.2626
- American Hospital Association. (2026). 2026 Health Care Workforce Scan. aha.org/aha-workforce-scan
Showing all 33 references.
AI Will Not Fix Healthcare’s Broken Workflows. It Will Industrialize Them. An interactive edition of the executive research report by Kelly Emrick, DHSc, PhD, MBA, BSRT(ARRT)R. Evidence current through August 20, 2026.
All calculators run entirely in your browser. No input is transmitted, stored, or shared. The E-R-A-A model, the Industrialized Dysfunction Risk model, and the AI Workflow Readiness Index are original proposals offered for organizational testing and are not validated instruments. Figures marked author-derived are arithmetic on the report’s published values and are labeled wherever they appear.
