Skip to content
The research library
Staffing levels and outcomesModerate evidence

Do 12-Hour Nursing Shifts Hurt Patient Care?

Nurses on 12-hour shifts rated safety and quality worse, but the outcomes are nurse ratings rather than patient records, and overtime tracked worse still.

Reviewed against primary sources on July 25, 2026 by the Soon operations research team. How we vet the evidence

The evidence in one line

In a cross-sectional survey of 31,627 nurses on adult medical and surgical wards in 488 European hospitals, working a shift of 12 hours or more was associated with 1.41 times the odds of grading patient safety on the unit as poor or failing, compared with shifts of 8 hours or less (OR 1.41, 95% CI 1.13 to 1.76; Griffiths et al., 2014). The design is correlational and every outcome is the nurse's own rating rather than a patient record, so this is a link between long shifts and worse nurse-reported safety, not evidence that patients were harmed. Working overtime on the shift carried a larger association than shift length did (OR 1.67, 95% CI 1.51 to 1.86).

What 31,627 nurses actually reported

The RN4CAST survey collected questionnaires from nurses on adult general medical and surgical wards in 488 hospitals across 12 European countries between June 2009 and June 2010. Of 54,140 questionnaires distributed, responses came back from 33,659 registered nurses (62%), and 31,627 of those worked the wards of interest and entered the analysis (Griffiths et al., 2014). Mean age was 38, 92% were female, 65% worked full time, and 76% reported on a day shift. Intensive care, high dependency, and long-term care units were deliberately excluded because their staffing and shift patterns differ, so none of these numbers transfer to ICU nursing or to any non-healthcare setting.

Half the sample had worked a shift of 8 hours or less (50%, n=15,930), 32% worked 8.1 to 10 hours (n=9,963), 4% worked 10.1 to 11.9 hours (n=1,160), 14% worked 12 to 13 hours (n=4,314), and 1% worked more than 13 hours (n=260). Those five counts account for all 31,627 nurses, with the percentages rounded. Shift length here means the hours worked on the most recent shift, not a contracted rotation. The survey did not capture usual pattern, fixed versus rotating work, weekly hours, break availability, or rest between shifts, all of which the authors name as unmeasured confounders. Total hours worked could not be measured directly either, so full-time versus part-time status stood in as a proxy.

Every outcome in both papers is a nurse's own rating: quality of care on the unit, a patient safety grade, and a count of care activities left undone on the last shift. There is no mortality, infection, fall, or medication error data anywhere in the study. The authors concede the outcomes are self-report, note that the clinical importance of the differences observed is unclear, and recommend that future work use objective measures. Nurse-reported quality and safety have been validated against mortality and failure to rescue elsewhere, but that is a separate body of work rather than a finding of this survey.

The 12-hour association is narrower than it looks

Against a reference group of shifts of 8 hours or less, nurses whose last shift ran 12 hours or more had 1.41 times the odds of grading patient safety poor or failing (95% CI 1.13 to 1.76) and 1.30 times the odds of rating quality of care poor or fair (95% CI 1.10 to 1.53). Those are odds, not case counts. An odds ratio of 1.41 means 41% higher odds of that rating, not 41% more safety incidents. The estimates come from multilevel models adjusted for shift type, ward type, patients per nurse, nurse age, full-time status, hospital size, high-technology status, and teaching status (Griffiths et al., 2014).

A dose-response gradient is largely absent for the patient-facing ratings. Odds of adverse quality and safety were raised at every shift length above 8 hours, but only the 12-hours-or-more category reached statistical significance, and the increase at 8.1 to 10 hours was described as marginal. A clean curve from 8 through 10 to 12 hours for quality and safety ratings is not supported by these estimates.

The one patient-facing outcome with a consistent gradient is care left undone. Nurses on shifts of 12 hours or more reported a 13% higher rate of care activities left undone on a 0 to 13 item count (RR 1.13, 95% CI 1.09 to 1.16), and that outcome reached significance across all shift lengths above 8 hours. This is a rate ratio on a self-reported count of omitted tasks, which makes it the closest thing in the study to a concrete operational measure and the most defensible one to replicate locally.

Overtime was the larger and more common signal

Working overtime on the shift, modeled independently of shift length, carried a stronger association with every outcome. Nurses who worked overtime had 1.67 times the odds of grading safety poor or failing (95% CI 1.51 to 1.86), 1.32 times the odds of rating quality poor or fair (95% CI 1.23 to 1.42), and a 29% higher rate of care left undone (RR 1.29, 95% CI 1.27 to 1.31). Every one of those estimates is larger than its shift-length counterpart (Griffiths et al., 2014).

Overtime was also more common than long shifts in this sample. A total of 27% of nurses (8,606) worked overtime on their last shift, ranging from 50% in England down to 12% in Poland, and from 0% to 80% across individual hospitals. The authors tested an interaction between shift length and overtime and it was not statistically significant, so in this data the two exposures appear to act independently rather than compounding one another.

The caveat that has to travel with the overtime numbers is that the survey could not distinguish mandatory from voluntary overtime, or paid from unpaid. That matters operationally, because a unit where nurses stay late by choice and a unit where they are held over are not the same problem even though they produce the same survey response. The very wide hospital range, from no overtime at all to 80% of nurses, shows how much the exposure varies between local settings.

The wider literature does not agree on where the worst point sits

The same cohort was analyzed for nurse-side outcomes. Compared with shifts of 8 hours or less, shifts of 12 hours or more were associated with higher odds of emotional exhaustion (aOR 1.26, 95% CI 1.09 to 1.46), low personal accomplishment (aOR 1.39, 95% CI 1.20 to 1.62), job dissatisfaction (aOR 1.40, 95% CI 1.20 to 1.62), and intention to leave the job due to dissatisfaction (aOR 1.29, 95% CI 1.12 to 1.48) (Dall'Ora et al., 2015). Two estimates sit on the edge of significance and should not be presented as settled: depersonalization (aOR 1.21, 95% CI 1.01 to 1.47) and dissatisfaction with work schedule flexibility (aOR 1.15, 95% CI 1.00 to 1.35). Job dissatisfaction did show a gradient below 12 hours, at aOR 1.15 (95% CI 1.05 to 1.25) for 8.1 to 10 hours and aOR 1.31 (95% CI 1.09 to 1.57) for 10.1 to 11.9 hours.

The wider literature is directionally consistent on long shifts but does not agree on where the worst point sits, and the authors say so themselves. They note prior US evidence that does not fit a simple linear hours effect: in one US study the odds of adverse reports of quality and safety were greater for nurses working 10 to 11 hours than for those working 12 hours or more, and a pediatric subsample showed elevated reports only above 13 hours. Both of those results reach this page through the RN4CAST discussion rather than through the primary US papers, so treat them as a reason the question stays open rather than as a settled US position. An earlier systematic review found insufficient evidence on shift length and nurse satisfaction, and nothing in the European survey rules out a mid-range peak.

Shift length in this survey is also heavily confounded with country, which is one reason the two literatures can point in different directions. Twelve-hour shifts were the norm in Poland (99% of nurses) and Ireland (79%), mixed in England (36% overall, 32% of day shifts and 37% of night shifts), and reported by fewer than 5% of nurses in Belgium, Germany, Greece, the Netherlands, Norway, and Sweden. The authors did not test interactions between country and shift work, so they can only estimate an average effect across all countries and cannot explore differences between them. Fieldwork also ran from June 2009 to June 2010, so this is evidence from a European survey now roughly 16 years old, predating post-2020 workforce conditions.

What this means for your schedule

  • Audit shift overrun before reopening the roster debate, because overtime carried a larger association with nurse-rated safety (OR 1.67) than shifts of 12 hours or more did (OR 1.41).
  • Measure care left undone on your own units, since it was the only patient-facing outcome with a consistent gradient across all shift lengths above 8 hours.
  • Separate mandatory from voluntary overtime in your own records, a distinction the survey could not make and one that changes which fix applies.
  • Do not carry these estimates into intensive care, high dependency, long-term care, or any non-nursing roster, because those settings were excluded by design.
  • Pair any shift-length review with retention data, since job dissatisfaction was higher at every shift length above 8 hours in the same cohort.

The business case

The strongest operational reading of this evidence is about shift overrun rather than roster design: nurses who worked overtime on their last shift had 1.67 times the odds of grading safety poor or failing, and 27% of the sample had worked overtime (Griffiths et al., 2014).

In the same cohort, shifts of 12 hours or more were associated with 1.40 times the odds of job dissatisfaction and 1.29 times the odds of intending to leave the job (Dall'Ora et al., 2015), which places shift length inside the retention conversation as well as the quality one.

Because the design is cross-sectional and every outcome is a nurse rating rather than a patient record, treat these as signals worth measuring on your own units, not as a forecast of what a roster change will deliver.

Frequently asked questions

Do 12-hour shifts cause worse patient outcomes?
Griffiths et al. (2014) measured what nurses thought, not what happened to patients, so the design cannot settle causation. In that cross-sectional survey of 31,627 European nurses, working a shift of 12 hours or more was associated with 1.41 times the odds of grading unit safety poor or failing (95% CI 1.13 to 1.76), and the authors state that causality cannot be inferred from the design. Every outcome was a nurse's own rating of quality, safety, or care left undone on the last shift, with no mortality, infection, fall, or medication error data. The finding is a link to nurse perception rather than to measured patient harm.
Why did overtime show a larger association than 12-hour shifts among European nurses?
In Griffiths et al. (2014), overtime on the last shift was modeled independently of shift length and tracked worse on every outcome: 1.67 times the odds of grading safety poor or failing (95% CI 1.51 to 1.86) against 1.41 for shifts of 12 hours or more, and a 29% higher rate of care left undone against 13%. Overtime was also more common, worked by 27% of nurses, and the tested interaction between the two exposures was not statistically significant, so in this data they appear to act independently rather than compounding.
Is a 10-hour shift a safe middle ground?
Griffiths et al. (2014) found the odds of adverse quality and safety ratings raised at every shift length above 8 hours among 31,627 European nurses, but only shifts of 12 hours or more reached statistical significance (OR 1.41 for a poor or failing safety grade), and the increase at 8.1 to 10 hours was described as marginal. The mid-range was estimated rather than skipped, but its raised odds did not reach significance, so it is an unresolved band rather than a demonstrated safe zone. The gradient below 12 hours did hold for care left undone, and for job dissatisfaction in Dall'Ora et al. (2015), which reported aOR 1.15 at 8.1 to 10 hours and aOR 1.31 at 10.1 to 11.9 hours.
Why does the RN4CAST survey not settle the 12-hour shift debate?
Because the literature disagrees and the data are old. Griffiths et al. (2014) point to US evidence where nurses' odds of adverse quality and safety reports were greater at 10 to 11 hours than at 12 hours or more, and a pediatric subsample where reports rose only above 13 hours. Fieldwork for the European survey of 31,627 nurses ran from June 2009 to June 2010, and 12-hour shifts were near universal in Poland (99%) and Ireland (79%) but rare in most other countries, confounding shift length with national health system.

Sources

Every figure on this page is drawn from a cited primary source and checked against the original publication.

  1. Griffiths, P., Dall'Ora, C., Simon, M., Ball, J., Lindqvist, R., Rafferty, A. M., Schoonhoven, L., Tishelman, C., Aiken, L. H., & RN4CAST Consortium (2014). Nurses' shift length and overtime working in 12 European countries: The association with perceived quality of care and patient safety. Medical Care, 52(11), 975โ€“981. https://doi.org/10.1097/MLR.0000000000000233

    Design: Cross-sectional survey of 31,627 registered nurses on adult medical and surgical wards in 488 hospitals across 12 European countries (June 2009 to June 2010), analyzed with multilevel generalized linear mixed models

  2. Dall'Ora, C., Griffiths, P., Ball, J., Simon, M., & Aiken, L. H. (2015). Association of 12 h shifts and nurses' job satisfaction, burnout and intention to leave: findings from a cross-sectional study of 12 European countries. BMJ Open, 5(9), e008331. https://doi.org/10.1136/bmjopen-2015-008331

    Design: Cross-sectional survey of the same cohort, 31,627 registered nurses in 2,170 general medical and surgical units within 488 hospitals across 12 European countries, nurse-reported outcomes

Cite these sources: BibTeX RIS

How independent are these sources?

Every source behind this article comes from one research program, RN4CAST Consortium, so the sources share investigators and methods. Agreement within a single program is weaker evidence than agreement between independent teams, because a shared measurement or modeling decision would produce the same pattern across all of them.

Why this page is graded moderate evidence

A consistent systematic review or meta-analysis at a lower grade, or a large observational study whose authors disclaim causality.

Who reviewed this

Every article in this library is checked against its primary sources by the Soon operations research team: each figure is traced back to the study it came from, and the wording is checked against the study design before publication. What that review covers

None of the studies cited here evaluated Soon.They examine scheduling practices, shift patterns, and working hours as studied by independent researchers, so their findings describe what those practices are associated with, not what any particular software produces.

This article summarizes published research for scheduling and operations decisions. It is not medical advice. Individual health questions belong with a qualified clinician.

Your next schedule could take 2 minutes.

Import your team, set your rules, hit auto-fill. Most teams are live the same day.

Try Soon free

30 days free ยท No credit card required

Already have an account?Sign in