Does a Structured Handover Actually Reduce Errors?
Error rates fell 23% after a nine-hospital handoff program, and the study design stops short of proving the program caused it.
Reviewed against primary sources on July 25, 2026 by the Soon operations research team. How we vet the evidence
The evidence in one line
Across 10,740 patient admissions at nine hospitals, the medical-error rate fell 23% after a resident handoff program was implemented, from 24.5 to 18.8 errors per 100 admissions (P<0.001), and preventable adverse events fell 30%, from 4.7 to 3.3 per 100 admissions (P<0.001) (Starmer et al., 2014). The study was a before-and-after intervention with no randomization and no concurrent control group, so its own authors describe the program as associated with those reductions rather than shown to have caused them. Six of the nine sites showed significant error reductions and three did not.
The headline result, and the number that keeps it honest
In 10,740 patient admissions across nine hospitals, the medical-error rate was 24.5 per 100 admissions before the handoff program and 18.8 per 100 admissions after it, a relative reduction of 23% (P<0.001) (Starmer et al., 2014). Preventable adverse events, the harms a better handoff would plausibly head off, went from 4.7 to 3.3 events per 100 admissions, a 30% relative reduction (P<0.001).
Read the scale carefully before quoting either figure. These are rates per 100 admissions, so a single admission can contribute more than one event, and neither number is the share of patients who came to harm. The 23% is a relative change between two observed rates, not an absolute risk difference and not a probability that any given handover goes wrong.
The most informative number in the paper is the one that did not move. Nonpreventable adverse events, which no communication change should be expected to touch, stayed flat at 3.0 and 2.8 events per 100 admissions (P=0.79). That internal control is what separates this study from a plain before-and-after chart, because a broad measurement drift or a general safety trend across the study period would likely have shifted both categories together.
Why this stops short of cause
The design is a prospective multicenter pre-post intervention study: nine hospitals, resident physicians, active surveillance for errors, and no randomization or concurrent control group. Secular improvement in patient safety over the study period, observation effects on the clinicians being watched, and quality initiatives running alongside the program cannot be ruled out. The authors' own conclusion says implementation of the program was associated with reductions in medical errors and in preventable adverse events, and that verb choice is deliberate.
Four things shipped at once: a mnemonic to standardize oral and written handoffs, handoff and communication training, a faculty development and observation program, and a sustainability campaign. Nothing in the study separates their contributions, so a claim that a standardized template alone produced the result is not supported. Anyone planning to adopt one component and expect the published effect is reading more into the paper than it reports.
Results also varied by site. Significant error reductions appeared at six of nine sites, which means three sites in a coordinated multicenter program with shared training and shared measurement did not show one. The published abstract reports rates and p-values rather than confidence intervals, so the precision around the 23% figure is not something this review can state, and no interval should be attached to it.
Handoff content changed, and the clock did not
The program moved the behavior it targeted. Across sites, significant increases were observed in the inclusion of all prespecified key elements in written documents and in oral communication during handoff, nine written elements and five oral elements, with P<0.001 for all 14 comparisons (Starmer et al., 2014). That is a coherent mechanism story: the content of handovers measurably improved in the same period the error rates fell.
The common operational objection is that structure costs time at shift change. This study did not find it. Oral handoff duration was 2.4 and 2.5 minutes per patient before and after (P=0.55), with no significant changes in resident workflow, including patient-family contact and computer time. Time-motion observation, not self-report, produced those workflow figures, which is a stronger basis than asking clinicians whether the new format felt slower.
One research group, one evaluation, one pediatric setting
One label belongs on this result before it travels anywhere. The senior author on Starmer et al. (2014) is Landrigan, the same investigator behind the resident duty-hour and extended-shift research cited elsewhere in this library, so the handoff finding and the duty-hour findings come from one Boston research group rather than independent teams. Shared surveillance procedures, shared method choices, and a shared working definition of what counts as a preventable error travel across all of that work.
This page also rests on a single pre-post evaluation. Later implementations of the same handoff bundle have been published, and none of them were checked for this review, so nothing here should be read as replication by a separate research group. Read the 23% as one program's result in nine hospitals during one study period, and read the six-of-nine site split as the best available warning that the same package does not land the same way everywhere.
The setting is narrow in ways that matter to anyone reading across from it. The population is resident physicians in nine hospitals handing over pediatric inpatients, every rate is expressed per 100 admissions, and each handover has a named receiver who assumes responsibility for a specific patient, a written record that already existed, and active surveillance able to find an error after the fact. A handover where the thing passed on is a queue, an open ticket, or a machine state, with no single receiver and no way to detect afterward what went wrong, is a different exchange wearing the same name.
The paper was published in 2014 with data collected earlier, before EHR-integrated handoff tools became common, so the written-handoff baseline it improved on may not match what a unit runs today. It remains the landmark citation in this literature, which is a reason to quote it accurately rather than a reason to treat its effect size as a forecast for a modern implementation.
What this means for your schedule
- Quote the 23% figure as a before-and-after observation from one evaluation in nine pediatric hospitals, never as a return you can expect from adding a handover template.
- Adopt the full bundle, standardized oral and written format plus training plus faculty observation plus a sustained campaign, because the study cannot tell you which piece carried the effect.
- Audit handover content against a prespecified element list before and after any change, since element inclusion is the thing this study showed measurably moved.
- Time your handovers on both sides of the change; oral handoffs ran 2.4 and 2.5 minutes per patient in the study, so a large time increase in your unit signals an implementation problem rather than the method.
- Plan for uneven results and investigate the units that do not improve, because three of nine sites in a coordinated program showed no significant error reduction.
The business case
The operational headline is that better handover content did not cost measurable time: oral handoffs ran 2.4 and 2.5 minutes per patient before and after (P=0.55), with no significant change in resident workflow (Starmer et al., 2014).
The spend is in training, faculty observation, and sustained attention rather than in longer shift overlaps, so this is a change-management program rather than a staffing cost.
Hold the expected effect loosely. Six of nine sites showed significant error reductions, the study had no control group, and no separate team's replication of the program is reviewed here, so results varied even inside one coordinated multicenter effort.
Frequently asked questions
- Did the handoff program cause the drop in medical errors?
- Causation stays open because the nine hospitals ran no comparison group. Starmer et al. (2014) recorded medical errors falling from 24.5 to 18.8 per 100 admissions after a resident handoff program was implemented, and the authors state that implementation was associated with reductions in medical errors and preventable adverse events rather than shown to have caused them. Secular safety trends over the study period and observation effects on the residents being watched cannot be ruled out.
- What does the 23% handoff-program reduction mean in absolute terms?
- In Starmer et al. (2014), the medical-error rate went from 24.5 to 18.8 per 100 admissions across 10,740 admissions (P<0.001), so 23% is the relative change between those two rates. Preventable adverse events fell from 4.7 to 3.3 per 100 admissions, a 30% relative reduction.
- Does a structured handover take longer to deliver?
- Time-motion observation found the clock essentially unchanged. Starmer et al. (2014) measured oral handoff duration in nine hospitals at 2.4 minutes per patient before the resident handoff program and 2.5 minutes after (P=0.55), and found no significant changes in resident workflow, including patient-family contact and computer time.
- Would the same handover result show up outside a hospital?
- Starmer et al. (2014) studied resident physicians handing over named pediatric inpatients, where the receiving clinician formally takes over a specific patient, a written record already exists, and active surveillance can find an error after the fact. Outside that setting the receiver is often nobody in particular, the work passed on is a queue rather than a patient, and no comparable audit exists to show whether error rates moved the way these did from 24.5 to 18.8 per 100 admissions. What transfers is the method: name the elements a handover must contain, then check whether they are said out loud, which is the behavior this program measurably changed.
Sources
Every figure on this page is drawn from a cited primary source and checked against the original publication.
Starmer, A. J., Spector, N. D., Srivastava, R., West, D. C., Rosenbluth, G., Allen, A. D., Noble, E. L., Tse, L. L., Dalal, A. K., Keohane, C. A., Lipsitz, S. R., Rothschild, J. M., Wien, M. F., Yoon, C. S., Zigmont, K. R., Wilson, K. M., O'Toole, J. K., Solan, L. G., Aylor, M., Bismilla, Z., ... Landrigan, C. P., & the I-PASS Study Group (2014). Changes in Medical Errors after Implementation of a Handoff Program. New England Journal of Medicine, 371(19), 1803โ1812. https://doi.org/10.1056/NEJMsa1405556
Design: Prospective multicenter pre-post intervention study in nine hospitals, uncontrolled quasi-experiment with no randomization and no concurrent control group, outcomes by active surveillance
Cite these sources: BibTeX RIS
Why this page is graded moderate evidence
A consistent systematic review or meta-analysis at a lower grade, or a large observational study whose authors disclaim causality.
Who reviewed this
Every article in this library is checked against its primary sources by the Soon operations research team: each figure is traced back to the study it came from, and the wording is checked against the study design before publication. What that review covers
None of the studies cited here evaluated Soon.They examine scheduling practices, shift patterns, and working hours as studied by independent researchers, so their findings describe what those practices are associated with, not what any particular software produces.
This article summarizes published research for scheduling and operations decisions. It is not medical advice. Individual health questions belong with a qualified clinician.
Keep reading
Is hospital care actually worse on the weekend?
Weekend admission is associated with higher mortality in English hospitals, but the largest study to test the staffing explanation did not find it.
Read →Why Do the Doctor Duty-Hour Trials Disagree?
Three randomized trials reported across five papers reach different conclusions, and the reason is which outcome each one measured.
Read →Does nurse staffing affect patient outcomes?
A four-article pillar on nurse workload and patient outcomes: what the landmark studies measured, what the California mandate research settled, and where the causal line actually sits.
Read →