Does an extra hour of work produce an extra hour of output?
Two panel studies, one from 1916 munitions plants and one from a modern call center, disagree about whether there is a threshold at all.
Reviewed against primary sources on July 25, 2026 by the Soon operations research team. How we vet the evidence
The evidence in one line
It depends on where the week sits, and the two best panel studies do not fully agree. Pencavel (2015) found that in British munitions plants, weeks below 49 hours showed output moving in proportion to hours (elasticity 1.002, SE 0.087), while weeks at 49 hours or more showed an elasticity of 0.172 (SE 0.089) that is not clearly distinguishable from zero. A modern call center panel found no threshold at all: within the same agent, 1 percent more daily hours went with about 0.9 percent more calls handled (Collewet & Sauermann, 2017), and since neither study randomized hours, both describe associations rather than proven causes.
Inside the 122 munitions weeks and the 33,123 agent-days
Pencavel (2015) returned to production records gathered by the British Health of Munition Workers Committee, mostly from 1916, and assembled 122 weekly observations on four groups of workers: 100 women in moderately heavy work, 40 women in light work, 56 men in heavy work, and 15 youths in light work. Weekly hours averaged 51.1 (SD 10.0) and ranged from 24.0 to 72.5, a span of scheduled hours no modern employer would run. He fitted least squares models with group fixed effects, including quadratic, cubic, and spline specifications.
Pooling all 122 observations in a log-log specification with group fixed effects gives an elasticity of output with respect to hours of 0.794 (SE 0.048). Splitting the sample at 49 weekly hours changes the picture entirely. For the 47 observations below 49 hours (mean 42.0 hours, mean output 5,461.1 units) the elasticity is 1.002 (SE 0.087), statistically indistinguishable from one. For the 75 observations at 49 hours or more (mean 56.7 hours, mean output 6,892.3 units) it falls to 0.172 (SE 0.089), which carries a t statistic of about 1.93 and therefore falls short of the 5 percent level. The honest reading of the upper range is that output rose at a decreasing rate, not that it rose by a precise 0.172 percent for every 1 percent of extra hours.
The spline models tell the same story in levels. With a knot at 49 hours, the marginal product of an hour is constant below the knot and declines above it, and the fitted output maximum sits at about 63 hours, with an R-squared of 0.736. Pencavel's own summary notes that fitted output at 70 hours is close to fitted output at 56 hours. Those statements describe a curve fitted to 122 observations, and the sparse upper tail of the hours distribution is exactly where the 63-hour maximum lives.
Collewet and Sauermann (2017) studied a very different setting: 33,123 agent-day records from 332 agents in one Dutch call center, running from mid-2008 to the first week of 2010. This is a part-time workforce, with mean daily working hours of 4.621 (SD 1.516), a maximum of 7.747, and mean weekly hours of 17.712. Their productivity measure is log inverse average handling time, in other words speed per call, against a mean handling time of 321.613 seconds and 55.795 calls answered per day.
With agent fixed effects, team fixed effects, hour-of-day dummies, day fixed effects, and a linear trend, the coefficient on log working hours is -0.113 (SE 0.013, p < 0.01, R-squared 0.385). Because the dependent variable is speed per call, the implied output elasticity is 1 minus 0.113, which the authors round to 0.9: within the same agent, days with 1 percent more working hours went with roughly 0.9 percent more calls handled. The weekly-hours model, estimated across 8,641 agent-weeks, gives a smaller coefficient of -0.078 (SE 0.031, p < 0.05), which is consistent with overnight recovery between days.
Munitions weeks kink at 49 hours, call center days never do
The two papers do not describe the same shape. Pencavel's data have a kink in them: below 49 hours an extra scheduled hour went with a full extra hour's worth of output, and above 49 hours additional time bought progressively less. Collewet and Sauermann tested for structural breaks and report finding none, so in their data the mild decline in speed per call is present across the observed range rather than switching on at any threshold.
Part of the disagreement is range. The Dutch agents never come near Pencavel's threshold: the longest observed day is 7.747 hours and the longest week is 43.046 hours, so their entire sample sits inside the region where the munitions data show proportional returns. Two readings survive that fact. Either fatigue accumulates smoothly from early in the day and Pencavel's below-threshold proportionality reflects a coarser weekly measure, or the two jobs differ enough that neither curve is a general law. Neither of these two papers separates those explanations.
Pencavel is careful about how far his own threshold travels. He selected the 49-hour knot partly on a priori grounds, since 48 hours became the post-war international standard, reports that other knots were tried and that 49 fit best, and states that the threshold applied to the munition workers he studied while "for other workers it may be more" (Pencavel, 2015). The knot was not discovered atheoretically from the data.
The two sets of numbers must not be spliced together. A claim such as "an elasticity of 0.9 above 48 hours" appears in neither paper. The 0.9 comes from a part-time sample with no threshold in it, and Pencavel's above-threshold elasticity is 0.172 (SE 0.089). Merging them produces a figure no author estimated and neither dataset supports.
The rest day is the sturdiest finding here
Across every regression Pencavel fitted, the coefficient on Sunday work is negative and statistically significant. In the spline specification that includes it, the Sunday coefficient is -858.162 weekly units (SE 125.727). Pencavel's own summary puts the output loss from denying workers a day of rest at about ten percent, a figure he draws from his fitted hours models rather than from dividing this one coefficient by an average week, so the two should not be read as the same calculation.
That specification also fits the data best of the set, with an R-squared of 0.812 against 0.736 for the same spline without the Sunday term, and it moves the fitted output maximum out to 68 hours. The practical reading is that when these plants ran seven days, weekly output was lower than the hours alone would predict, and this pattern is more consistent across specifications than anything in the upper tail of the hours curve. It remains 1916 factory data with no control group, so it is a strong association rather than a demonstrated cause.
Wartime shells, part-time call days, and nothing in between
Neither study randomized anything. Pencavel has group fixed effects on historical records, with no instrument, no discontinuity, and no control group, so his results are associations between weeks with more hours and weeks with more output. Collewet and Sauermann have a stronger design, because daily hours in their call center were set by central scheduling against expected demand rather than by worker preference, and they do use the word effect. Even so, there is no randomization, instrument, or discontinuity, so the claim their data support is that when the same agent worked a longer day, average handling time was longer.
The populations are narrow and nothing like each other. One is British munitions workers, most of them women, doing repetitive physical work in wartime plants observed mostly in 1916. The other is 332 part-time agents in a single Dutch call center between 2008 and 2010, averaging 4.621 hours a day. Neither paper studied salaried knowledge work, and neither claims its numbers transfer to it. Applying either curve to an office, a design team, or an engineering group is an extrapolation the reader owns, not a finding either author published.
The call center result is sensitive to how hours are counted. The daily hours-available-for-calls specification gives -0.078 (SE 0.015, p < 0.01), a point estimate that happens to match the weekly-hours model quoted earlier while coming from a different specification with its own standard error, so the two should never be cited as one result. Measured instead as hours present at the workplace, the coefficient falls to -0.016 (SE 0.015) and loses statistical significance, because slack time absorbs the association. Excluding shifts starting before 09:00 or ending after 21:00 moves the coefficient only to -0.105 (p < 0.01), so timing is not driving it. Every coefficient quoted here is estimated on average handling time, which measures speed per call rather than how well the call was resolved.
Collewet and Sauermann note that the productivity decrease is mild in a mostly part-time sample and suggest fatigue effects would be stronger at full-time hours, but that is their conjecture, not an estimate they produced. Pencavel names the small number of observations as a shortcoming of his data. Neither paper reports confidence intervals, only standard errors, and neither speaks to remote work, hybrid work, or four-day-week trials, since the data end in 1916 and 2010 respectively.
What this means for your schedule
- Treat the 49-hour threshold as a fact about Pencavel's munition workers rather than a general limit, since he states the number may be higher for other workers.
- Never quote a merged figure such as an elasticity of 0.9 above 48 hours, because that number appears in neither paper and neither dataset supports it.
- Protect the weekly rest day before arguing about the length of the week, since the Sunday coefficient was negative and significant in every regression Pencavel fitted, an output loss he puts at about ten percent.
- Define hours precisely before drawing any conclusion, because the call center association shrank from -0.113 to -0.016 and lost significance when hours meant time present at the workplace rather than time worked.
The business case
The commercial point in this literature is a ceiling, not a target. In Pencavel's (2015) munitions data, output above 49 weekly hours rose at a decreasing rate with a fitted maximum near 63 hours, and in a modern call center panel the same agent's speed per call slipped as the day lengthened (Collewet & Sauermann, 2017), so scheduled hours and delivered output are not interchangeable planning units.
Neither study priced anything, neither randomized hours, and neither examined salaried knowledge work, so no cost figure travels from these papers to your budget. What does travel is a planning distinction: in Pencavel's (2015) models the largest single negative coefficient sat on Sunday work, -858.162 weekly units (SE 125.727), rather than on the marginal hour of an already long weekday.
Frequently asked questions
- Does an extra hour of scheduled work add a full hour of output?
- It depended on where the week sat. In Pencavel's (2015) analysis of 122 weekly observations from British munitions plants, weeks below 49 hours had an output-hours elasticity of 1.002 (SE 0.087), consistent with proportional returns, while weeks at 49 hours or more had an elasticity of 0.172 (SE 0.089), a t statistic of about 1.93 that is not clearly distinguishable from zero. These are associations drawn from historical production records, not the result of an experiment.
- Is 48 hours a universal productivity threshold?
- No. Pencavel (2015) placed the knot at 49 hours partly on a priori grounds, because 48 hours became the post-war international standard, and he states that the threshold applied to the munition workers he studied while other workers may differ. Collewet and Sauermann (2017) tested for structural breaks in their call center panel and report finding none, so no threshold appears in their data at all.
- Did the call center study prove that long hours cause lower productivity?
- It did not establish causation. Collewet and Sauermann (2017) estimate a coefficient of -0.113 (SE 0.013, p < 0.01) on log daily hours in a within-agent model with team, day, and hour-of-day controls, implying an output elasticity of about 0.9, and their identification rests on hours being set by central scheduling rather than by worker choice. There is no randomization, instrument, or discontinuity, and the sample averages 4.621 hours a day, so the finding describes part-time days in one workplace.
- What does the evidence say about working through the weekly rest day?
- Pencavel (2015) found the Sunday-work coefficient negative and statistically significant in every regression he fitted, at -858.162 weekly units (SE 125.727) in the spline specification that includes it, and he summarizes the output loss from denying workers a day of rest as about ten percent. The estimate comes from 1916 British munitions plants with no control group and no randomized comparison, so it describes weeks in those wartime factories, and no study reviewed here retests it.
Sources
Every figure on this page is drawn from a cited primary source and checked against the original publication.
Pencavel, J. (2015). The productivity of working hours. The Economic Journal, 125(589), 2052โ2076. https://doi.org/10.1111/ecoj.12166
Design: Retrospective panel of 122 weekly observations on four groups of British munitions workers, mostly 1916, analyzed by least squares with group fixed effects and hours splines; no randomization, instrument, or control group
Collewet, M., & Sauermann, J. (2017). Working hours and productivity. Labour Economics, 4796โ106. https://doi.org/10.1016/j.labeco.2017.03.006
Design: Within-worker panel of 33,123 agent-day records from 332 part-time agents in one Dutch call center, with agent, team, and day fixed effects; hours set by central scheduling rather than randomized
Cite these sources: BibTeX RIS
Why this page is graded moderate evidence
A consistent systematic review or meta-analysis at a lower grade, or a large observational study whose authors disclaim causality.
Who reviewed this
Every article in this library is checked against its primary sources by the Soon operations research team: each figure is traced back to the study it came from, and the wording is checked against the study design before publication. What that review covers
None of the studies cited here evaluated Soon.They examine scheduling practices, shift patterns, and working hours as studied by independent researchers, so their findings describe what those practices are associated with, not what any particular software produces.
This article summarizes published research for scheduling and operations decisions. It is not medical advice. Individual health questions belong with a qualified clinician.
Keep reading
Does Overtime Cause Injuries, or Do Hazardous Jobs Use More of It?
The largest US cohort found long hours associated with more injuries even after adjusting for industry and occupation, while the meta-analysis that followed is far less settled.
Read →What Did the Four-Day Week Trials Actually Measure?
Six months, 141 organizations, and a burnout drop the authors themselves say is probably too large.
Read →Does better scheduling actually pay?
One randomized experiment, two large surveys, and an honest account of how much of this field we could not verify.
Read →