Skip to content
The research library
Scheduling and business performanceStrong evidence

What Happened When a Retailer Randomized Stable Scheduling?

A randomized controlled field experiment in scheduling research: what it measured, what it earned, and where it stops.

Reviewed against primary sources on July 19, 2026 by the Soon operations research team

The evidence in one line

Between November 2015 and August 2016, Gap Inc. let researchers randomize a bundle of scheduling practices across 28 of its stores, and the intent-to-treat analysis found that responsible scheduling practices increased store productivity by 5.1%, a result of increasing sales by 3.3% and decreasing labor by 1.8% (Kesavan et al., 2022). Randomization is what earns the word increased here, and no other study in this pillar has it. The result is also narrower than the headline suggests, because the control stores were already receiving two-week advance notice and no on-call shifts, so 5.1% is the marginal return on additional practices layered on that base.

What was actually randomized

The experiment ran as a randomized controlled field experiment across 28 stores in the San Francisco and Chicago metros, 19 assigned to treatment and 9 to control, from November 2015 to August 2016 (Kesavan et al., 2022). It covered roughly 150,000 shifts of about 1,500 employees and was analyzed intent-to-treat with difference-in-differences, meaning stores were compared by the group they were assigned to rather than by how faithfully they carried the changes out. Stores that adopted the bundle loosely are averaged in with stores that adopted it well, which makes the headline number the estimate to quote for the experiment under its actual implementation conditions.

The treatment was a package, not a practice. Treatment stores received fixed shift times, consistent employee-to-shift assignment, core scheduling, part-time-plus with a 20-hour minimum, a shift-swap app, and targeted staffing (Kesavan et al., 2022). Every one of those levers moved at once, on purpose, because the researchers were testing whether a coherent scheduling regime pays rather than which knob is the profitable one. The design can show whether the regime paid, but it cannot isolate which practice did the work.

Which estimate to quote, and which one people quote instead

The peer-reviewed intent-to-treat estimate is a 5.1% gain in store productivity, decomposed into a 3.3% increase in sales and a 1.8% decrease in labor (Kesavan et al., 2022). It averages every assigned store, including those that implemented the bundle unevenly. The same analysis reported a 16.4% treatment-on-treated estimate adjusted for adherence, but that model-dependent estimate describes this experiment and is neither an expected upside nor an upper bound for another operation.

A separate practitioner report on the same experiment, by Williams et al. (2018), circulates far more widely than the journal article and was not peer reviewed. It reports productivity up 5%, equal to $6.20 more revenue per labor hour than control stores, and median sales up 7% in treatment stores. The 7% is a median across treatment stores, while the 3.3% is the peer-reviewed mean effect estimated intent-to-treat, so the figures cannot be swapped for one another or folded into the same productivity decomposition. When someone quotes a 7% sales lift from this experiment, they are quoting the practitioner report rather than the peer-reviewed finding.

The money, and why the dollar figures do not port

Williams et al., 2018 states the result in store-level dollars: treatment stores averaged $4,363 more per week than controls, and median sales in treatment stores rose 7%. Across 19 treatment stores over 35 weeks the report puts added revenue at an estimated $2.9 million, against an out-of-pocket intervention cost of about $31,200. Those are report figures rather than peer-reviewed estimates, and the ratio between them is the reason this experiment gets cited in board decks.

What the experiment did not test is the story people attach to those dollars. The intuitive explanation, that the same people worked the same shifts often enough to know the floor and the stock and each other, is interpretation rather than a tested finding, and it should be labeled as such when it goes in front of a finance team.

The dollar figures themselves do not travel. The $6.20 more revenue per labor hour is denominated in Gap's 2015-16 revenue base, in two high-cost metros, in an apparel format with its own price points and margin structure, so a hospital, a contact center, or a distribution center has no basis for expecting that number. What does travel is the shape of the finding: productivity moved because sales rose while labor hours fell, and that ratio is measurable against any operation's own baseline.

Five things that narrow what this proves

Start with the design. Of the 28 stores, 25 were randomized, with 3 pretest stores carried into treatment by design, so assignment was not perfectly clean. More consequentially, the control group was not a no-intervention baseline: two-week advance notice and the elimination of on-call shifts were rolled out chain-wide to treatment and control stores alike from October 1 2015. The estimate is therefore the marginal effect of five additional practices layered on that base, not stable scheduling measured against chaotic scheduling. An operation still running on-call shifts and one-week notice sits further from the control group than it looks, and this experiment does not price what closing that gap is worth.

The remaining limits govern attribution. Free targeted additional staffing hours, funded centrally, went to 13 of 19 treatment stores, which confounds schedule stability with extra labor input, so some share of the sales gain may have been bought rather than organized. The treatment was a bundle, so no single practice can be credited, and adopting fixed shift times alone gives no claim on a proportional slice of 5.1%. And the setting was one retailer, one apparel format, two metros, nine months, so external validity to healthcare, contact centers, or manufacturing is unestablished rather than merely untested in public.

What this means for your schedule

  • Quote 5.1% as the productivity effect and say intent-to-treat when you quote it, because that is the estimate peer review stands behind.
  • Check whether your current baseline already includes two-week advance notice and no on-call shifts before assuming this experiment describes your starting point.
  • Budget for the added staffing hours, since 13 of 19 treatment stores received them free from a central fund and your rollout will not.
  • Measure productivity as sales per labor hour against your own baseline, because the dollar figures here are denominated in Gap's 2015-16 apparel economics.
  • Refuse to credit any single practice you happen to favor, since the bundle was only ever tested whole.

The business case

This is the strongest causal evidence in this library, which makes it the one place a scheduling investment case can rest on caused rather than correlated with.

The peer-reviewed result was a 5.1% intent-to-treat productivity gain, from a 3.3% sales increase and a 1.8% decrease in labor. The same analysis reported a 16.4% adherence-adjusted estimate, but the two estimates answer different questions and do not define an expected range for another operation (Kesavan et al., 2022).

Fund it as a pilot with a real control group rather than a chain-wide rollout, because the conditions that limit this study, a single retailer and a bundled treatment with centrally funded extra hours, are precisely what your own pilot can resolve.

Frequently asked questions

Is the Gap scheduling study strong enough to base an investment on?
It is the strongest causal evidence in this library, because it randomized 19 treatment and 9 control stores rather than surveying workers after the fact (Kesavan et al., 2022). Randomization is what licenses causal language that survey evidence cannot support. It remains one retailer in two metros over nine months, so treat it as evidence that the effect is real somewhere, not as a forecast of its size in your operation.
Why do I see a 7% sales lift quoted instead of 3.3%?
Because those figures come from different documents and different estimators. The 3.3% is the peer-reviewed intent-to-treat mean effect on sales (Kesavan et al., 2022). The 7% is a median across treatment stores reported in the practitioner write-up of the same experiment (Williams et al., 2018), which was not peer reviewed. They are not interchangeable, and combining them yields a number that no analysis ever estimated.
Does 5.1% understate the effect?
Not necessarily. The 5.1% productivity gain reported by Kesavan et al., 2022 is the intent-to-treat estimate, so it averages every store assigned to treatment under the implementation conditions observed in the experiment. The same analysis reported a 16.4% treatment-on-treated estimate adjusted for adherence, but that model-dependent result is not a forecast for disciplined teams. Treat 5.1% as the observed estimate from this experiment and 16.4% as a conditional within-study estimate, not as a floor and reachable ceiling.
Which part of the bundle did the work?
The study cannot say, and neither can anyone citing it. Kesavan et al., 2022 randomized a bundle of six scheduling practices across 19 treatment stores and evaluated it as a single package, so no individual practice carries its own estimate and none can claim a proportional slice of the 5.1% productivity gain. Compounding the ambiguity, 13 of 19 treatment stores also received free targeted additional staffing hours funded centrally, so stability and extra labor input moved together.

Sources

Every figure on this page is drawn from a cited primary source and checked against the original publication.

  1. Kesavan, S., Lambert, S. J., Williams, J. C., & Pendem, P. K. (2022). Doing Well by Doing Good: Improving Retail Store Performance with Responsible Scheduling Practices at the Gap, Inc. Management Science, 68(11), 7818โ€“7836. https://doi.org/10.1287/mnsc.2021.4291

    Randomized controlled field experiment (28 stores, roughly 150,000 shifts, about 1,500 employees)

  2. Williams, J. C., Lambert, S. J., Kesavan, S., Fugiel, P. J., et al. (2018). Stable Scheduling Increases Productivity and Sales: The Stable Scheduling Study. [Report]. Center for WorkLife Law, UC Hastings College of the Law. https://worklifelaw.org/publications/Stable-Scheduling-Study-Report.pdf

    Practitioner report on the same randomized experiment (not peer reviewed)

None of the studies cited here evaluated Soon.They examine scheduling practices, shift patterns, and working hours as studied by independent researchers, so their findings describe what those practices are associated with, not what any particular software produces.

This article summarizes published research for scheduling and operations decisions. It is not medical advice. Individual health questions belong with a qualified clinician.

Your next schedule could take 2 minutes.

Import your team, set your rules, hit auto-fill. Most teams are live the same day.

Try Soon free

30 days free ยท No credit card required

Already have an account?Sign in