The Law of Averages

Healthcare professionals discussing an operational problem around a conference table

Interactive Infographic · Evidence Synthesis for Healthcare Leaders

Why Leaders Must Look Beyond the Average

Power laws, uncertainty, and evidence-based decisions in healthcare and business.

  • Focused narrative synthesis
  • Ten cited sources
  • Conceptual decision matrix

Figure 1. Examining assumptions before choosing an intervention. AI-generated editorial illustration; the people and setting do not depict an actual study or event.

Chapter 1

The same number can describe very different realities

Averages can conceal consequential differences in organizational performance. Three situations can look alike on an executive dashboard and still call for three different decisions.

Situation one

An acceptable average, a hidden wait

A department can report an acceptable average while a small group of patients waits far too long.

Situation two

Steady results, a single point of failure

An organization can show steady monthly performance while depending on a single vulnerable system.

Situation three

Repeated disappointment, then value

A new service can disappoint repeatedly before a different approach reveals substantial value.

See it: one average, many realities

Each dot is one of 200 simulated report turnaround times. Move the slider. The mean stays pinned at 24.0 hours while the experience behind it changes.

EvenlyVery unevenly

The interactive chart needs JavaScript. The values below show the default setting.

  • Report finished within 48 hours
  • Report beyond the 48 hour example threshold
  • Mean, fixed at 24.0 hours
  • Median
24.0 hMean (what the dashboard shows)
16.4 hMedian (the typical report)
50.0 h90th percentile
68.2 h95th percentile
193.7 hLongest single report
22 of 200Reports beyond 48 hours

The dashboard reports 24.0 hours. Half of all reports are finished within 16.4 hours, yet 22 of 200 take longer than 48 hours and the slowest takes 193.7 hours.

Illustrative teaching example built for this infographic from 200 simulated values rescaled to a fixed mean. It is not data from the report or from any facility, and it does not claim that turnaround times follow a particular distribution. The 48 hour threshold is an example only.

The practical recommendation

Match decisions to their consequences.

The approach can support innovation while retaining accountability for patient outcomes.

Preserve

dependable service.

Investigate

severe losses.

Test

uncertain opportunities with limited exposure.

The evidence at a glance

Select a figure to open the chapter that explains it.

Next: why a skewed graph is not automatically a power law.

Chapter 2

What the distribution can tell us

Normal, lognormal, and power-law patterns are distinct mathematical possibilities, not interchangeable labels for a skewed graph (Newman, 2005). Each one raises a different leadership question.

Three patterns, compared

Choose a pattern, then switch the view to see how each tail behaves on logarithmic axes.

The interactive chart needs JavaScript. The three patterns are described alongside.

Shape view: the normal curve is symmetric around its center, the lognormal curve leans right, and the power-law tail declines slowly beyond its threshold.

Conceptual comparison based on Newman (2005) and Clauset et al. (2009). The curves use illustrative parameters chosen for this infographic. No distribution is fitted to organizational data.

Normal

Statistical feature

Symmetric around a center; relatively thin tails.

In plain terms

Most observations sit near the center, and very distant observations become increasingly unlikely.

Leadership question

Is this model adequate for the process and time period?

Lognormal

Statistical feature

Positive and right-skewed; multiplicative effects can generate it.

In plain terms

It can emerge when positive quantities grow through multiplicative effects. It is asymmetric and can produce substantial inequality.

Leadership question

How different are typical outcomes from the mean?

Power-law tail

Statistical feature

Scale-related decline beyond a threshold; extreme outcomes can matter greatly.

In plain terms

The tail declines more slowly still. Increasing an outcome by a fixed factor reduces its exceedance probability by a predictable factor, over the range where the model holds.

Leadership question

How sensitive is the decision to rare events and upper limits?

Where the question came from

The article starts from a Veritasium video that contrasts familiar averages with distributions in which rare outcomes exert disproportionate influence. Its useful challenge is to examine the assumptions behind our expectations.

The scientific literature also supplies reasons to resist treating power laws as a universal explanation.

One finding, three separate questions

Take an executive who learns that a small share of accounts generates most revenue. Each step to the right needs its own evidence.

1. Concentration finding

A small share of accounts generates most revenue.

What it is: a description of where outcomes are concentrated. It can show where to look.

2. Statistical question

Does that pattern follow a power law?

What it needs: estimation, goodness-of-fit testing, and comparison against alternative distributions (Clauset et al., 2009).

3. Causal question

Will concentrating future resources on those accounts improve performance?

What it needs: evidence that the action works, judged against a credible comparison. Chapter 4 shows why.

Claim check

Four statements leaders often hear about power laws. Decide whether each is accurate, then compare your answer with the research.

“Every power law makes averages meaningless.”

InaccurateSome idealized, unbounded power laws have no finite population mean or variance, depending on their exponent. Others have both. With finite real-world limits, moments may be finite even when extreme observations dominate practical decisions (Newman, 2005).

“A roughly straight line on logarithmic axes proves a power law.”

InaccurateIt is insufficient evidence. Clauset, Shalizi, and Newman (2009) combine estimation with goodness-of-fit testing and comparison against alternative distributions.

“Strong scale-free structure is the norm in real networks.”

InaccurateBroido and Clauset (2019) found that strong evidence of scale-free structure was uncommon in a large collection of empirical networks, and that lognormal models often fitted as well or better. The finding concerns network connectivity, not every organizational outcome, but it cautions against assuming universality.

“Leaders should ask whether their estimates are stable and what the available sample misses.”

AccurateLeaders must ask whether their estimates are stable, what the available sample misses, and which consequences an average obscures (Newman, 2005).

Next: why a success story is weaker evidence than it looks.

Chapter 3

What extreme success can and cannot prove

Large differences in outcomes do not reveal how much of the difference comes from each of five sources.

  • Quality
  • Effort
  • Timing
  • Accumulated advantage
  • Chance
Controlled experiment

Popularity is not an independent measure of quality

14,341

participants in an artificial music market (Salganik, Dodds, and Watts, 2006)

When other participants’ choices were made visible, two things increased.

Inequality of outcomes
Unpredictability of success

Quality still mattered, but social influence affected which products became successful.

Where the evidence stops

This was a constructed cultural market. It does not establish the same mechanism in every industry.

Management inference

Unequal opportunity can look like unequal potential

The highly visible project

May attract additional support because it already appears promising.

  • Sponsorship
  • Staff
  • Attention

The less visible project

May never receive comparable support.

  • Sponsorship
  • Staff
  • Attention

If the final results are evaluated without considering these differences, leaders can mistake unequal opportunity for unequal underlying potential.

Where the evidence stops

This is a management inference from the research, not a claim that the experiment measured employee performance.

Survivorship bias: add the failures back in

Two simulated groups of 30 organizations with the same true average performance. One follows a steady practice, the other a risky one. Switch the view to see what a leader observes when the failures drop out of sight.

The interactive chart needs JavaScript. Summary: among survivors only, the risky practice appears to outperform (58.0 against 50.0). With every organization counted, both groups average 50.0.

  • Steady practice
  • Risky practice
  • Did not survive
50.0Average you observe, steady practice
58.0Average you observe, risky practice
22 of 30Risky organizations still visible

Among survivors, the risky practice looks better (58.0 against 50.0), and 9 of the 10 best performers followed it. Nothing about the practice improved: the weakest results simply left the sample.

Illustrative simulation of the argument in Denrell (2003): observing surviving organizations while overlooking failures can make risky practices appear beneficial even when they have no favorable relationship with performance in the full population. Values are constructed for this infographic and are not study data.

The review standard

A credible decision review needs the unsuccessful attempts as well as the celebrated results.

The successful executive who persisted through repeated setbacks is memorable. The organizations that persisted in equally expensive mistakes are less likely to become leadership case studies.

Physical model

An analogy should raise questions, not settle them

Bak, Tang, and Wiesenfeld (1987), self-organized criticality

What it offers

The model showed how interactions in a dynamical system can generate complex behavior. It is a useful analogy for how small disturbances might propagate.

Where the evidence stops

It does not establish that a hospital, career, or business is governed by that model. Similar-looking distributions can arise through different mechanisms.

Organizational learning

The unit of analysis is the decision, not the organization

March (1991), exploration and exploitation

Exploit

Existing operations need reliable execution.

Example: strict medication verification

Both, at once

Explore

Uncertain opportunities need room for learning.

Example: testing appointment outreach

Two words worth redefining

Persistence

Continuing to search and learn while retaining the ability to stop a specific unsuccessful idea.

Not repeating an unchanged intervention indefinitely.

Consistency

Remains valuable in experimentation.

Comparable measurements, clear definitions, and dependable follow-up are essential to learning whether anything improved.

Next: a randomized trial that tested the argument in healthcare.

Chapter 4

A healthcare test of the argument

Concentrated need can help identify a population deserving attention, but it does not demonstrate that a proposed intervention will work.

The Camden care-transition trial

Finkelstein and colleagues (2020) randomly assigned medically and socially complex patients to the program or to usual care.

800patients randomly assigned
62.3% vs 61.7%readmitted within 180 days, program vs usual care
0.82 pointsadjusted difference; 95% CI −5.97 to 7.61

The chart needs JavaScript. The three values above are the trial results cited in the report.

The interval spans zero. The trial did not demonstrate reduced readmissions for that intervention and population.

Values as cited in the report from Finkelstein et al. (2020), New England Journal of Medicine. The difference is in percentage points, program minus usual care.

What the trial does not show

It does not show that complex patients cannot benefit from support, or that all care-management programs fail.

Uncertainty remains

The finding concerns one program and one population. The interval runs from 5.97 points lower to 7.61 points higher, so the estimate is compatible with no effect.

What the trial does show

Identifying an extreme group and demonstrating an effective response are different tasks.

Why a before-and-after comparison can overstate benefit

Selecting patients during unusually high utilization can produce an apparent later improvement even without the intervention. This is regression to the mean.

The schematic needs JavaScript. In words: utilization falls after selection in both the program group and the comparison group, so the fall alone cannot be credited to the program.

Seen alone, the program group appears to improve sharply after enrollment. It is tempting to credit the program.

Schematic illustration of the concept described in the report. The lines are not trial data and carry no numeric scale.

Four randomized trials

A constructive direction: disciplined revision

759 firms

Camuffo and colleagues (2024)

Training in a scientific approach to decisions increased termination of ideas and encouraged a limited number of strategic changes.

What it supports

Explicit hypotheses and disciplined revision.

Where the evidence stops

The trials took place in entrepreneurial settings. They do not establish a universal formula for commercial success, and application to clinical operations remains an extrapolation.

For healthcare leaders

Make the proposed mechanism visible

If an access initiative is intended to reduce missed appointments:

  1. 1Specify how it changes the barriers patients experience.
  2. 2Measure whether that change actually occurred.
  3. 3Check whether attendance improved relative to a credible comparison.

A plausible explanation deserves a test; an attractive distribution does not complete it.

Selected evidence and its limits

Source and designContributionBoundary
Clauset et al. (2009)
Statistical methods and applications
Tests power-law fit and competing explanations.A plausible fit alone does not identify a causal mechanism.
Salganik et al. (2006)
Controlled music-market experiment
Social influence changed inequality and predictability.Artificial market; not a workplace intervention.
Finkelstein et al. (2020)
Randomized healthcare trial
No demonstrated reduction in 180-day readmissions.One program and population; uncertainty remains.
Camuffo et al. (2024)
Four randomized trials of 759 firms
Scientific decision training increased idea termination and focused strategic changes.Entrepreneurial settings; transfer to healthcare requires testing.

Table 2 of the report. Author synthesis of the cited publications. The studies answer different questions and are not pooled into a common effect estimate.

Next: a way to match the intensity of evaluation to what is at stake.

Chapter 5

Matching the decision to its consequences

The framework separates the potential to expand a benefit from the consequences of failure. Those two dimensions guide the intensity of evaluation and the limits placed on exposure.

Place a decision on the matrix

Answer the two questions, or select a quadrant directly.

Potential to expand benefits
Consequences if the idea fails

Potential to expand benefits

Limited

Substantial

Consequences if the idea fails

Containable

Severe

Limited benefit potential · Containable consequences

Improve routine work

  • Standardize a reliable process.
  • Measure variation and burden.
  • Retain changes that help.

Existing operations need reliable execution. Comparable measurements, clear definitions, and dependable follow-up show whether anything improved.

Substantial benefit potential · Containable consequences

Run a bounded experiment

  • State a testable hypothesis.
  • Limit exposure and set stop rules.
  • Expand after credible evidence.

Uncertain opportunities need room for learning. A favorable early result should earn a stronger test. It should not automatically trigger organization-wide adoption.

Limited benefit potential · Severe consequences

Protect essential service

  • Prioritize continuity and safeguards.
  • Test recovery before disruption.
  • Maintain a workable fallback.

An average cannot say how long essential work can continue. Scenario testing and recovery exercises can reveal vulnerabilities even when major incidents are too few to estimate a precise tail.

Substantial benefit potential · Severe consequences

Redesign before expansion

  • Reduce potential harm first.
  • Use staged evidence and oversight.
  • Do not trade safety for a large payoff.

The aim is to support innovation while retaining accountability for patient outcomes.

Try an example from the report

The examples are placed to illustrate the reasoning of the report. Your own assessment of a similar decision may differ.

Original conceptual synthesis by Kelly Emrick. Qualitative guidance, not a validated scoring instrument. The matrix has not been empirically validated.

Before the experiment begins: build a one-page brief

The report lists what to settle before testing an uncertain idea. Fill in each element, then print or copy the brief.

No elements defined yet. An experiment without a stated comparison and stop rule is difficult to learn from.

These are proposed operating practices rather than a tested package of interventions. Entries are kept in this browser only so that you can return to them. Nothing is sent to a server.

Before you scale

A favorable early result should earn a stronger test

Consider whether the improvement depends on any of the following.

Replication across another setting can reveal these dependencies before the organization commits more resources.

Ten projects are not always ten independent opportunities

Switch each shared dependency on or off to see how many separate bets remain.

The interactive diagram needs JavaScript. With all three dependencies shared, the ten projects form a single connected group.

1Separate bets remaining
10 of 10Projects exposed to the same failure

All ten projects are linked through shared dependencies. Examine common dependencies when setting a combined exposure limit.

Illustrative diagram built for this infographic. Which projects share which dependency is invented to show the principle. A low-cost idea can become expensive if it creates irreversible obligations or shifts hidden work to clinicians and patients.

Next: the same reasoning applied to an imaging department.

Chapter 6

Applying the reasoning in radiology operations

Four situations in an illustrative imaging department. Each one asks a different question of the same leadership team.

A new appointment reminder process

The decision

An uncertain opportunity. The useful move is a test in a defined service, not a department-wide launch.

If random assignment is feasible and ethically appropriate, it strengthens causal interpretation. If it is not, the evaluation should explain what alternative comparison is used and which biases remain.

Design the test
  • Test the process in a defined service.
  • Compare attendance with an appropriate concurrent group.
  • Assess patient understanding.
  • Assess staff workload.
  • Assess access across relevant patient groups.

A prolonged information-system outage

What the average hides

An average uptime percentage cannot by itself answer the questions that matter during a long outage.

Scenario testing and recovery exercises can reveal vulnerabilities even when there are too few major incidents to estimate a precise tail distribution.

Ask instead
  • How long can essential work continue?
  • Which functions depend on the failed service?
  • Can staff execute the fallback process?

Report turnaround time

What the average hides

Measurement should retain the experiences hidden by aggregation. A facility-wide mean can conceal where the problems are.

Tail percentiles calculated from small samples should be interpreted cautiously. None of these practices requires claiming that turnaround times follow a power law.

Report alongside the mean
  • The median.
  • Upper percentiles.
  • The number of clinically important delays beyond a defined threshold.
  • Results segmented by urgency, service, and patient circumstances.

High-cost care

What concentration hides

High spending can reflect serious illness, necessary treatment, fragmented care, or several factors at once. A concentration analysis can indicate where to investigate.

Lower spending alone is not a sufficient patient-centered outcome.

Consider explicitly
  • Decisions about care should follow evidence about needs and effective interventions.
  • Access, function, experience, and safety may change in different directions.
  • Each of those four deserves explicit consideration.

Hypothetical application described in the report, not a claim about observed performance at a particular facility.

The first practical action for a leadership team

  1. 1Select one decision currently justified primarily by an average or a success story.
  2. 2Examine its distribution and the cases left out of the story.
  3. 3State the mechanism through which the proposed action is expected to help.
  4. 4Decide whether the immediate need is a controlled experiment, a more reliable routine, or protection against a severe disruption.
  5. 5Assign responsibility for reviewing the evidence on a specified date.

Two expensive errors this standard protects against

Stopping too soon

Abandoning useful possibilities because early results are modest.

Continuing too long

Sustaining ineffective programs because their original rationale sounded persuasive.

The leadership standard

An evidence-based leader should be able to explain both why an initiative is worth trying and what would change that judgment.

The value of understanding extreme outcomes lies in better decisions about where to investigate, where to experiment, and where dependable protection is essential.

Next: the sources, and what kind of evidence each one provides.

Chapter 7

Sources and method

The sources play different evidentiary roles. Filter the list to see which ones are statistical, experimental, theoretical, or science communication.

Methodological note

This article is a focused narrative synthesis, not a systematic review or meta-analysis. It combines statistical and theoretical work with selected experiments.

Its healthcare applications and decision matrix are recommendations for local evaluation, not estimates of proven intervention effects. The 2024 replication is included alongside foundational studies because the evidentiary roles of these sources differ.

Source distinction

The Veritasium video is the starting point of the article and a science-communication source. The other references are peer-reviewed publications.

The managerial recommendations are identified as synthesis or illustrative application. The interactive demonstrations in this infographic are illustrations built for teaching and are labeled where they appear.

References

  • Theory and modelsBak, P., Tang, C., & Wiesenfeld, K. (1987). Self-organized criticality: An explanation of the 1/f noise. Physical Review Letters, 59(4), 381-384. https://doi.org/10.1103/PhysRevLett.59.381
  • Statistical methodsBroido, A. D., & Clauset, A. (2019). Scale-free networks are rare. Nature Communications, 10, Article 1017. https://doi.org/10.1038/s41467-019-08746-5
  • Experiments and trialsCamuffo, A., Gambardella, A., Messinese, D., Novelli, E., Paolucci, E., & Spina, C. (2024). A scientific approach to entrepreneurial decision-making: Large-scale replication and extension. Strategic Management Journal, 45(6), 1209-1237. https://doi.org/10.1002/smj.3580
  • Statistical methodsClauset, A., Shalizi, C. R., & Newman, M. E. J. (2009). Power-law distributions in empirical data. SIAM Review, 51(4), 661-703. https://doi.org/10.1137/070710111
  • Theory and modelsDenrell, J. (2003). Vicarious learning, undersampling of failure, and the myths of management. Organization Science, 14(3), 227-243. https://doi.org/10.1287/orsc.14.2.227.15164
  • Experiments and trialsFinkelstein, A., Zhou, A., Taubman, S., & Doyle, J. (2020). Health care hotspotting: A randomized, controlled trial. New England Journal of Medicine, 382(2), 152-162. https://doi.org/10.1056/NEJMsa1906848
  • Theory and modelsMarch, J. G. (1991). Exploration and exploitation in organizational learning. Organization Science, 2(1), 71-87. https://doi.org/10.1287/orsc.2.1.71
  • Statistical methodsNewman, M. E. J. (2005). Power laws, Pareto distributions and Zipf’s law. Contemporary Physics, 46(5), 323-351. https://doi.org/10.1080/00107510500052444
  • Experiments and trialsSalganik, M. J., Dodds, P. S., & Watts, D. J. (2006). Experimental study of inequality and unpredictability in an artificial cultural market. Science, 311(5762), 854-856. https://doi.org/10.1126/science.1121066
  • Science communicationVeritasium. (2025, November 26). You’ve (likely) been playing the game of life wrong [Video]. YouTube. https://www.youtube.com/watch?v=HBluLfX2F_k

Showing 10 of 10 references.

Kelly Emrick, DHSc, PhD, MBA, BSRT(ARRT)R

Interactive companion to the article Why Leaders Must Look Beyond the Average, October 5, 2026.

The decision matrix is an original conceptual synthesis, not an empirically validated instrument. Interactive demonstrations are illustrative and are labeled where they appear.

Homekellyemrick.com