Skip to content
The research library
Scheduling and business performanceModerate evidence

Does better scheduling actually pay?

One randomized experiment, two large surveys, and an honest account of how much of this field we could not verify.

Reviewed against primary sources on July 19, 2026 by the Soon operations research team

The evidence in one line

The strongest evidence that scheduling practice moves business results comes from the one randomized field experiment on business outcomes we were able to verify: a bundle of responsible scheduling practices at Gap Inc. raised store labor productivity 5.1% on an intent-to-treat basis, from a 3.3% sales increase and a 1.8% reduction in labor hours (Kesavan, Lambert, Williams, & Pendem, 2022). Everything else in this pillar is survey evidence, which links unpredictable schedules to worse worker outcomes but cannot establish that changing a schedule produces them. That asymmetry, one experiment alongside a large body of correlational work, is the honest shape of the field.

The one experiment, and the caveats that travel with it

The strongest evidence in this pillar is a randomized field experiment in retail. Kesavan et al. (2022) worked with Gap Inc. across 28 stores and about 1,500 employees, assigning some stores to a bundle of responsible scheduling practices and leaving the others to operate as usual. The 5.1% productivity result is the intent-to-treat estimate, which averages every store assigned to the practices, including stores that implemented them unevenly. Because assignment was randomized, causal wording is defensible for this experiment in a way it is nowhere else on this pillar, but the estimate is not a guaranteed floor for another operation.

The caveats are part of the finding, not a footnote to it. This is one retailer in one sector at one moment, the sample is small by the standards of retail analytics, and the treatment was a bundle, so the experiment cannot tell you which practice inside it did the work. It also cannot tell you what the same bundle would do in a hospital, a warehouse, or a contact center, where the demand pattern and the labor model look nothing like a mall storefront. The dedicated article on the experiment walks through the design, and a second article covers its secondary result about who left the treatment stores, which is often more useful to an operations manager than the headline productivity figure.

The survey evidence, and what a survey can and cannot say

The second cluster comes from The Shift Project, and its value is scale and breadth rather than the ability to establish cause. Schneider and Harknett (2019) surveyed 27,792 hourly workers at 80 large firms between June 2016 and October 2017, and the questionnaire covers a range of practices instead of one, from on-call shifts to canceled shifts to short notice to start times that move week to week. That breadth makes it a usable map of how common each practice is, with 14% of those workers reporting a canceled shift and 26% reporting on-call shifts. A later study of 37,263 hourly workers at 127 large firms (Schneider & Harknett, 2021) reports a steep gradient between schedule unpredictability and material hardship, expressed as predicted probabilities that adjust for wages, income, and hours.

Both are cross-sectional surveys on a non-probability sample, so the correct verb is associated, never causes. The data capture workers at one point in time, which cannot separate the possibility that unstable schedules worsen worker outcomes from the possibility that workers already under strain end up in the least stable jobs. It also matters that the sharpest quantities in both papers are model outputs rather than observations: the 2019 paper simulates what removing on-call shifts would be associated with, and the 2021 paper reports adjusted predicted probabilities. Neither is a measured before-and-after. The article on schedule instability and worker well-being, and the article on the hardship gradient, each take one of these studies and work through what it supports.

What we verified, and what we could not

This library publishes a figure only when we can trace it to a primary source and check it against the original. Applying that standard across eight areas of scheduling research, two clusters cleared it: the Gap Inc. experiment and The Shift Project surveys. Six did not clear it in that pass: fair workweek law evaluations, nurse staffing ratios and mortality, schedule-control turnover experiments, mandatory overtime and error, burnout drivers, and demand-matched staffing.

That is a statement about our search, not a verdict on those literatures. Several of them are large and well regarded, and a different search, better database access, or more time would probably surface usable numbers in some. What we will not do is publish a figure we could not trace to a primary source, so those areas stay empty here until they can be filled properly. If you have been shown a scheduling return figure with no traceable citation behind it, this gap is the reason to treat it carefully.

How to weigh one experiment against tens of thousands of survey responses

The temptation is to average the two clusters into a single claim that better scheduling pays. Resist that, because they do not carry equal weight and they answer different questions. One randomized experiment that moved a business metric is stronger causal evidence than any volume of survey responses, while the surveys are far stronger on scale and on describing what unpredictability looks like in workers' lives. Stacking them into one list of benefits misrepresents both.

In practice, the Gap experiment is the reference point when someone asks whether scheduling changes can move store performance, and the Shift Project studies are the reference point when someone asks what unpredictability is associated with for the people living it. Keeping those two questions apart is most of the discipline this pillar asks for, because the moment they blur, a survey correlation starts doing work that only an experiment can do.

What this means for your schedule

  • Quote the 5.1% productivity result as evidence from one randomized experiment at one retailer, not as an industry benchmark you should expect to hit.
  • Use associational wording for the survey findings, since a cross-sectional survey cannot show that fixing a schedule improves worker outcomes.
  • Ask for the primary source behind any scheduling return figure you are shown, including the ones on this page.
  • Treat the missing evidence on fair workweek laws, nurse ratios, and burnout as an open question, not as proof those levers fail.
  • Send each question to the article that covers it rather than arguing from the category of better scheduling as a whole.

The business case

The strongest evidence that scheduling practice moves store performance found 5.1% higher labor productivity, and it came from both sides of the ledger: a 3.3% sales increase and a 1.8% reduction in labor hours (Kesavan et al., 2022).

That result covers 28 stores and about 1,500 employees at a single retailer, so read it as proof the mechanism can work rather than as a forecast for your own operation.

The rest of the published evidence describes worker outcomes rather than financial ones, so a business case built on it should be argued on retention and risk grounds instead of a return figure the literature does not supply.

Frequently asked questions

Is there real proof that better scheduling improves business results?
There is one strong piece of evidence: a randomized field experiment at Gap Inc. where a bundle of responsible scheduling practices raised store labor productivity 5.1% on an intent-to-treat basis (Kesavan et al., 2022). Intent-to-treat means the estimate averages every store assigned to the practices, including those that applied them unevenly. It covered 28 stores and about 1,500 employees at a single retailer, so it establishes what happened in that experiment without guaranteeing an effect or its size elsewhere. Outside that experiment, the verified published evidence on this pillar describes worker outcomes rather than business results.
Why is the survey evidence described as associational rather than causal?
Schneider and Harknett ran cross-sectional surveys on a non-probability sample, 27,792 hourly workers in the 2019 study and 37,263 in the 2021 study. That design measures workers at one moment rather than following them through a schedule change, so it can show that workers reporting less predictable schedules also report worse outcomes, but it cannot rule out the reverse, that people already under strain get sorted into the least stable jobs. The sharpest quantities in each paper are also model outputs rather than observations: the 2021 hardship gradient is a set of predicted probabilities adjusting for wages, income, and hours, while the 2019 paper reports descriptive group comparisons alongside simulations of what removing a practice would be associated with.
Does the missing evidence mean fair workweek laws or nurse staffing ratios do not work?
No. Six of the eight scheduling research areas we searched, fair workweek law evaluations and nurse staffing ratios and mortality among them, produced nothing we could confirm against primary sources in that pass. That is a limit of our search rather than proof those literatures are weak, and several of them are large and well regarded. It does mean you should ask for the underlying study whenever someone quotes a number from one of those areas, because on this pillar the Gap Inc. experiment (Kesavan et al., 2022) is the only source that gave us a business-outcome figure we could trace.
How should a manager act when the evidence base is this thin?
Lean on the one randomized result, the 5.1% productivity gain at Gap Inc. (Kesavan et al., 2022), only where a retail-style operation resembles it, and treat the Shift Project survey findings (Schneider & Harknett, 2019, 2021) as description rather than forecast. Those surveys cover 27,792 and 37,263 hourly workers, which makes them strong on what unpredictability looks like in workers' lives and silent on what changing a schedule would deliver. Where no verified evidence exists, say so and measure your own operation instead of borrowing an untraceable number. The methods article on this pillar shows how to check a study before you lean on it.

Sources

Every figure on this page is drawn from a cited primary source and checked against the original publication.

  1. Kesavan, S., Lambert, S. J., Williams, J. C., & Pendem, P. K. (2022). Doing Well by Doing Good: Improving Retail Store Performance with Responsible Scheduling Practices at the Gap, Inc. Management Science, 68(11), 7818โ€“7836. https://doi.org/10.1287/mnsc.2021.4291

    Randomized controlled field experiment (28 stores, roughly 150,000 shifts, about 1,500 employees)

  2. Schneider, D., & Harknett, K. (2019). Consequences of Routine Work-Schedule Instability for Worker Health and Well-Being. American Sociological Review, 84(1), 82โ€“114. https://doi.org/10.1177/0003122418823184

    Cross-sectional survey (27,792 hourly workers at 80 large firms, 2016-2017, non-probability sample)

  3. Schneider, D., & Harknett, K. (2021). Hard Times: Routine Schedule Unpredictability and Material Hardship among Service Sector Workers. Social Forces, 99(4), 1682โ€“1709. https://doi.org/10.1093/sf/soaa079

    Cross-sectional survey (37,263 hourly workers at 127 large firms, 2017-2019, non-probability sample)

None of the studies cited here evaluated Soon.They examine scheduling practices, shift patterns, and working hours as studied by independent researchers, so their findings describe what those practices are associated with, not what any particular software produces.

This article summarizes published research for scheduling and operations decisions. It is not medical advice. Individual health questions belong with a qualified clinician.

Your next schedule could take 2 minutes.

Import your team, set your rules, hit auto-fill. Most teams are live the same day.

Try Soon free

30 days free ยท No credit card required

Already have an account?Sign in