Why Do the Doctor Duty-Hour Trials Disagree?
Three randomized trials reported across five papers reach different conclusions, and the reason is which outcome each one measured.
Reviewed against primary sources on July 25, 2026 by the Soon operations research team. How we vet the evidence
The evidence in one line
In a randomized crossover trial in intensive care units, interns made 35.9 percent more serious medical errors while working a traditional schedule with extended shifts of 24 hours or more than while working a schedule capped at 16 hours, 136.0 versus 100.1 per 1,000 patient-days (P<0.001) (Landrigan et al., 2004). Two later cluster-randomized trials, one in general surgery and one in internal medicine, found that giving programs flexibility over shift-length limits was noninferior on their primary patient outcomes against prespecified margins (Bilimoria et al., 2016; Silber et al., 2019). Those are different endpoints rather than a refutation: the crossover trial counted directly observed errors on 2,203 patient-days, while the later trials counted death and complications across tens of thousands of patients, where most errors leave no trace.
Three trials, five papers, one survey
What gets quoted as four or five studies is three randomized trials reported across five papers, plus one observational survey. Landrigan et al. (2004) and Lockley et al. (2004) report two outcome sets from a single crossover trial at one academic institution. Bilimoria et al. (2016) reports the FIRST trial in general surgery. Silber et al. (2019) and Basner et al. (2019) report patient safety and sleep from the single iCOMPARE trial in internal medicine. Barger et al. (2005) is a nationwide survey, not a trial, so it can describe association only.
The crossover trial has the strongest internal validity and the narrowest reach. Twenty interns worked two three-week intensive care rotations, one under each schedule, so each intern served as their own control (Lockley et al., 2004). Seventeen of the twenty worked more than 80 hours per week on the traditional schedule, with a mean of 84.9 hours and a range of 74.2 to 92.1, while all of them stayed under 80 hours on the capped schedule, mean 65.4 and range 57.6 to 76.3. Interns worked 19.5 fewer hours per week (P<0.001), slept 5.8 more hours per week (P<0.001), and had less than half the rate of attentional failures during on-call nights (P=0.02).
What the cap changed inside the ICU
Across 2,203 patient-days involving 634 admissions, interns made 35.9 percent more serious medical errors on the traditional schedule than on the intervention schedule, 136.0 versus 100.1 per 1,000 patient-days (P<0.001) (Landrigan et al., 2004). Serious errors that were not intercepted were 56.6 percent more frequent under the traditional schedule (P<0.001), and total serious errors on the critical care units were 22.0 percent higher, 193.2 versus 158.4 per 1,000 patient-days (P<0.001).
The breakdown matters more than the headline. Serious medication errors were 20.8 percent more frequent under the traditional schedule, 99.7 versus 82.5 per 1,000 patient-days (P=0.03), while serious diagnostic errors were 5.6 times as many, 18.6 versus 3.3 per 1,000 patient-days (P<0.001). The outcome that depends most on sustained attention is where the gap was widest. The abstract reports P values and no confidence intervals, so no interval belongs on either figure.
Randomization and blinded outcome raters make this a causal result inside its setting, and the setting is small: intensive care units at a single academic institution, with interns who could not be blinded to their own schedules. Errors were identified by a four-pronged method including continuous direct observation, which detects far more than a registry abstraction or a claims file will, and which also means these rates cannot be lined up against rates produced by a different method.
Why FIRST and iCOMPARE read as a contradiction
FIRST randomized 117 US general surgery residency programs to standard duty-hour rules or a flexible policy, and analyzed 138,691 patients. The 30-day rate of postoperative death or serious complications was 9.1 percent under the flexible policy and 9.0 percent under the standard policy (P=0.92), with an unadjusted odds ratio of 0.96 and a 92 percent confidence interval of 0.87 to 1.06 (P=0.44), satisfying the trial's noninferiority criteria (Bilimoria et al., 2016). Quote that interval at 92 percent, a width tied to the one-sided noninferiority test, rather than converting it to a 95 percent interval.
The resident-reported results are the part usually left out. Among 4,330 residents, dissatisfaction with overall education quality was 11.0 percent under flexible policies and 10.7 percent under standard ones (P=0.86), and dissatisfaction with well-being was 14.9 percent versus 12.0 percent (P=0.10). Residents under flexible policies were less likely to report leaving during an operation, 7.0 percent versus 13.2 percent (P<0.001), or handing off active patient issues, 32.0 percent versus 46.3 percent (P<0.001). Flexibility bought continuity at the bedside, and the well-being comparison did not reach significance in either direction.
iCOMPARE randomized 63 internal-medicine residency programs and tested the change in unadjusted 30-day mortality from the pretrial year to the trial year, comparing that change between arms rather than comparing the arms head to head. Flexible programs were at 12.5 percent in the trial year against 12.6 percent in the pretrial year, and standard programs were at 12.2 percent against 12.7 percent (Silber et al., 2019). Read that as mortality in flexible programs not worsening by more than the prespecified 1 percentage point relative to standard ones: the point estimate has flexible programs improving 0.4 points less, and the one-sided 95 percent upper limit of that difference is 0.93 percent (P=0.03).
Two of the trial's other prespecified tests did not clear the bar. Differences in changes for 7-day readmission, patient safety indicators, and Medicare payments were all below 1 percentage point, but the noninferiority criterion was not met for 30-day readmissions or prolonged length of hospital stay (Silber et al., 2019). Noninferiority also asks whether a policy is not meaningfully worse against a stated margin, never whether it is equal or better, so quoting the mortality result alone overstates how reassuring this trial was.
Sleep cleared its margin, alertness did not
The iCOMPARE substudy tracked 205 interns at six flexible programs and 193 interns at six standard programs with 14 days of actigraphy. Sleep averaged 6.85 hours per 24 hours under flexible policies (95% CI 6.61 to 7.10) and 7.03 hours under standard ones (95% CI 6.78 to 7.27), a between-group difference of -0.17 hours with a one-sided lower confidence limit of -0.45 hours against a margin of -0.5 hours (P=0.02), which met noninferiority (Basner et al., 2019). Subjective sleepiness on the Karolinska Sleepiness Scale was also noninferior, differing by 0.12 points with an upper limit of 0.31 points against a 1-point margin (P<0.001).
Alertness did not clear the same bar. On the brief Psychomotor Vigilance Test the between-group difference was -0.3 lapses, but the upper limit of the one-sided 95 percent confidence interval reached 1.6 lapses against a noninferiority margin of 1 lapse (P=0.10), so noninferiority was not established (Basner et al., 2019). Lapse counts were 5.3 under flexible policies (95% CI 3.7 to 7.0) and 5.7 under standard ones (95% CI 4.1 to 7.3). A test that fails to establish noninferiority is not evidence of harm, and it is also not permission to claim equivalence.
One exposure in this literature was never randomized by anybody. Barger et al. (2005) surveyed 2,737 interns who filed 17,003 monthly reports, and reporting a motor vehicle crash after an extended shift carried 2.3 times the odds of reporting one after a non-extended shift (95% CI 1.6 to 3.3). The same intern reported both the shift and the crash, and nobody assigned who drove home after which shift, so that is an association between self-reports, and a survey sits below three randomized trials in this page's evidence order.
None of these results describe current practice either. The crossover trial ran before the 2003 and 2011 ACGME reforms, so its traditional schedule no longer exists in US training, and FIRST and iCOMPARE both predate the 2017 revision that restored 24-hour shifts for interns. The populations are not interchangeable: intensive care interns at one institution, general surgery residents nationally, and internal medicine residents nationally, across different specialties, eras, and outcome definitions. Landrigan, Lockley, and Barger come from one research group with a shared senior author, and Silber and Basner come from the iCOMPARE group, so this is three research programs rather than five separate confirmations. Nothing here speaks to non-physician or non-US shift workers.
What this means for your schedule
- Name the outcome you care about before citing a duty-hour trial, because directly observed errors and claims-based mortality behaved differently across these studies.
- Read every noninferiority result as not meaningfully worse against a stated margin, and check which prespecified tests failed before quoting the reassuring one.
- Treat the 16-hour cap evidence as evidence about intern errors in intensive care, not as a proven mortality intervention in any other setting.
- Keep the drive home inside any extended-shift policy, but weigh it as survey evidence rather than trial evidence, because none of these three trials randomized the commute.
- Track handover frequency whenever you change shift limits, because FIRST residents under flexible policies reported handing off active patient issues far less often.
The business case
Duty-hour policy is not one lever with one outcome. A 16-hour cap cut directly observed error rates in a randomized crossover trial, while flexible rules stayed inside a prespecified noninferiority margin in two much larger trials, one measuring death or serious complications abstracted from a surgical quality registry and the other 30-day mortality in Medicare claims.
Decide which endpoint your organization will be judged on before you change a schedule, and accept that a registry or claims evaluation across tens of thousands of patients cannot detect most of what continuous direct observation can.
Noninferior is not the same as equally safe, so any policy defended with the phrase no harm found should be checked against the prespecified outcomes these trials did not pass.
Frequently asked questions
- How many randomized trials tested capping doctors' shifts?
- Three, reported across five papers. Landrigan et al. (2004) and Lockley et al. (2004) report one randomized crossover trial in intensive care units, Bilimoria et al. (2016) reports the FIRST trial in general surgery, and Silber et al. (2019) and Basner et al. (2019) report the iCOMPARE trial in internal medicine. Barger et al. (2005) is a prospective survey rather than a trial, so it supports association only.
- Does noninferiority mean flexible duty hours were as safe?
- No. FIRST measured 30-day death or serious complications in 138,691 surgical patients and iCOMPARE measured 30-day mortality in Medicare claims, and each asked only whether flexibility was not meaningfully worse than standard rules against a prespecified margin (Bilimoria et al., 2016; Silber et al., 2019). iCOMPARE cleared its mortality margin with a one-sided 95 percent upper limit of 0.93 percentage points against a margin of 1 percentage point, and did not meet its noninferiority criterion for 30-day readmissions or prolonged length of hospital stay.
- How much did the 16-hour cap change interns' hours and sleep?
- In the crossover trial, 17 of 20 interns worked more than 80 hours per week on the traditional schedule, mean 84.9 hours, and all of them stayed under 80 hours on the capped schedule, mean 65.4 hours. They worked 19.5 fewer hours per week and slept 5.8 more hours per week, both P<0.001 (Lockley et al., 2004).
- Did flexible duty hours change interns' sleep or alertness?
- In the iCOMPARE actigraphy substudy, sleep averaged 6.85 hours per 24 hours under flexible policies and 7.03 hours under standard ones, which met the prespecified noninferiority margin of -0.5 hours (Basner et al., 2019). Alertness did not clear its bar: on the brief Psychomotor Vigilance Test the upper confidence limit reached 1.6 lapses against a 1-lapse margin (P=0.10), so noninferiority was not established, which is neither a finding of harm nor grounds for claiming equivalence.
Sources
Every figure on this page is drawn from a cited primary source and checked against the original publication.
Landrigan, C. P., Rothschild, J. M., Cronin, J. W., Kaushal, R., Burdick, E., Katz, J. T., Lilly, C. M., Stone, P. H., Lockley, S. W., Bates, D. W., & Czeisler, C. A. (2004). Effect of reducing interns' work hours on serious medical errors in intensive care units. New England Journal of Medicine, 351(18), 1838โ1848. https://doi.org/10.1056/NEJMoa041406
Design: Prospective randomized within-subject crossover trial in intensive care units at one academic institution, errors rated by physicians blinded to schedule
Lockley, S. W., Cronin, J. W., Evans, E. E., Cade, B. E., Lee, C. J., Landrigan, C. P., Rothschild, J. M., Katz, J. T., Lilly, C. M., Stone, P. H., Aeschbach, D., & Czeisler, C. A. (2004). Effect of reducing interns' weekly work hours on sleep and attentional failures. New England Journal of Medicine, 351(18), 1829โ1837. https://doi.org/10.1056/NEJMoa041404
Design: Hours, sleep, and attention outcomes from the same randomized crossover trial, 20 interns across two three-week intensive care rotations
Barger, L. K., Cade, B. E., Ayas, N. T., Cronin, J. W., Rosner, B., Speizer, F. E., & Czeisler, C. A. (2005). Extended work shifts and the risk of motor vehicle crashes among interns. New England Journal of Medicine, 352(2), 125โ134. https://doi.org/10.1056/NEJMoa041401
Design: Prospective nationwide web-based cohort survey (2,737 US first-year residents, 17,003 monthly self-reports of work hours, crashes, near-misses, and involuntary sleep episodes)
Bilimoria, K. Y., Chung, J. W., Hedges, L. V., Dahlke, A. R., Love, R., Cohen, M. E., Hoyt, D. B., Yang, A. D., Tarpley, J. L., Mellinger, J. D., Mahvi, D. M., Kelz, R. R., Ko, C. Y., Odell, D. D., Stulberg, J. J., & Lewis, F. R. (2016). National cluster-randomized trial of duty-hour flexibility in surgical training. New England Journal of Medicine, 374(8), 713โ727. https://doi.org/10.1056/NEJMoa1515724
Design: National pragmatic cluster-randomized noninferiority trial of 117 US general surgery residency programs in 2014-2015, analyzing 138,691 patients whose outcomes were abstracted from a surgical quality registry
Silber, J. H., Bellini, L. M., Shea, J. A., Desai, S. V., Dinges, D. F., Basner, M., et al. (2019). Patient safety outcomes under flexible and standard resident duty-hour rules. New England Journal of Medicine, 380(10), 905โ914. https://doi.org/10.1056/NEJMoa1810642
Design: Cluster-randomized noninferiority trial of 63 US internal-medicine residency programs in 2015-2016, difference-in-difference analysis of Medicare claims
Basner, M., Asch, D. A., Shea, J. A., Bellini, L. M., Carlin, M., Ecker, A. J., et al. (2019). Sleep and alertness in a duty-hour flexibility trial in internal medicine. New England Journal of Medicine, 380(10), 915โ923. https://doi.org/10.1056/NEJMoa1810641
Design: Actigraphy and alertness substudy of the iCOMPARE cluster-randomized trial, 398 interns at 12 programs over 14 days
Cite these sources: BibTeX RIS
Why this page is graded strong evidence
A randomized trial, or a finding that an umbrella review or meta-analysis graded at its top tier after pooling many underlying studies.
Who reviewed this
Every article in this library is checked against its primary sources by the Soon operations research team: each figure is traced back to the study it came from, and the wording is checked against the study design before publication. What that review covers
None of the studies cited here evaluated Soon.They examine scheduling practices, shift patterns, and working hours as studied by independent researchers, so their findings describe what those practices are associated with, not what any particular software produces.
This article summarizes published research for scheduling and operations decisions. It is not medical advice. Individual health questions belong with a qualified clinician.
Keep reading
Is clinician burnout linked to patient safety?
The most cited paper on this question was retracted in 2020, and the meta-analysis that replaced it comes from substantially the same research group.
Read →Do 12-Hour Nursing Shifts Hurt Patient Care?
Nurses on 12-hour shifts rated safety and quality worse, but the outcomes are nurse ratings rather than patient records, and overtime tracked worse still.
Read →Does nurse staffing affect patient outcomes?
A four-article pillar on nurse workload and patient outcomes: what the landmark studies measured, what the California mandate research settled, and where the causal line actually sits.
Read →