The Firefighting Capability Trap and Hero Engineers
The Firefighting Capability Trap and Hero Engineers
What we actually think: "hero engineer" is practitioner slang, but the system that manufactures
heroes has a formal name and thirty years of study. The capability trap is two competing uses of
the same hour — work now versus capability later — wired into a reinforcing loop: pressure pulls
hours from improvement into firefighting, eroded capability manufactures more fires after a
delay, and the new fires justify more firefighting (Repenning & Sterman 2001, 2002). The loop is
independently documented in manufacturing (Du Pont), product development (auto/motorcycle),
university facilities (Lyneis & Sterman 2016, >85% of maintenance hours reactive), and software
(Rahmandad & Repenning 2016's "adaptation trap": local workload search systematically overloads
the org). We hold the mechanism at "moderate" — convergent models plus field studies, but
small-N, no controlled trials.
The trap's most useful formal property is bistability. In Repenning's multi-project model a
temporary ~20% workload shock is absorbed and the system recovers; ~25% tips it past an unstable
threshold into a firefighting equilibrium from which "absent an additional intervention, it
never recovers" — and the tipping point's location is set by steady-state utilization. We
verified the 20/25% simulation numbers and the 60%→70% / 40%→25% one-year map in the MIT-hosted
working papers, and we hold the whole thing at "suggestive" because these are model
illustrations, not measured constants — a caveat the authors make themselves and that Nordhaus's
"measurement without data" critique of the tradition sharpens. Escape is equally structural:
every documented exit pays a worse-before-better cost on purpose (BP Lima's maintenance spend
ballooned 30% in six months before pump MTBF went from 12 to 58 months), and escapes fail when
the investment is too small or the transitional dip is read as failure.
Why do rational people choose the trap side? Because the credit function is rigged. Prevented
fires are counterfactuals nobody observes: Healy & Malhotra's electoral panel shows relief
spending rewarded and preparedness spending unrewarded despite ~1;
the action-bias literature (goalkeepers who jump although staying put is optimal) supplies the
psychology; and Perlow's ethnography shows the same machine inside a software team — rewards
tracked "managers' perceptions of the individual's heroics," the top-ranked engineers were
annotated "works 80-100 hours/week," and a $5,000 prevention request languished four months.
We cap the credit-asymmetry claim at "moderate" per the verified packs' honest gap: no direct
fixer-vs-preventer promotion experiment exists; the domains transfer to organizational
promotion only by analogy. The practitioner record then states the leading-indicator thesis
nearly verbatim — Google SRE's "Because heroism masks systemic problems, the systemic problems
are never fixed," and a team case where heroic paging response preceded the formal overload
declaration — but that is one company's doctrine, so it stays "anecdotal" no matter how quotable.
The bounds are load-bearing, not decorative. Slack is inverted-U for innovation (Nohria &
Gulati), so the prescription is protected improvement allocation, not organizational fat. Lean
is the sharpest apparent counterexample and actually the strongest confirmation: NUMMI ran
near-Toyota productivity on two days of inventory by deleting buffer slack while
institutionalizing improvement slack (kaizen teams, andon authority) — the distinction between
the two slacks, taken from Adler's chapter, is the essay's key conceptual move. HROs design
containment capacity deliberately because not everything is preventable; skilled response is not
the pathology — reward-dominant, chronic dependence on it is. And DORA's allocation gradient
(49/21 vs 38/27) is real but non-monotonic (medium performers report more rework than low) and
rests on self-reported, vendor-sponsored, cross-sectional surveys.
Limitations of this ingestion: the seed report cites via unrecoverable tokens, so all
bibliography was rebuilt from the five verified source packs and re-checked; 46 of 51 evidence
excerpts were primary-checked against MIT-hosted PDFs, sre.google, PMC, Cambridge Core,
DSpace, and the DORA PDFs. Discrepancies caught and corrected during re-verification: the Healy
& Malhotra 1 spent on preparedness is worth about 5,000
request was funded when the situation "became desperate" (the report's "became a crisis" is a
paraphrase); the SRE heroism quote begins with "Because"; the 200x/24x/3x DevOps multiples are
2016 figures, not 2018 (2018 is 46x/2,555x/2,604x/7x); and two corrected DOIs from the packs
were confirmed resolving (Patt & Zeckhauser 10.1023/A:1026517309871; Nordhaus 10.2307/2230846).
The report's quantitative synthesis (the dC/dt allocation model, the v_h > κ·v_p·e^{-rτ}
inequality, the hero-load indicator H(t)) is its own construct — each component loop is sourced,
but the assembled model is not encoded as a claim and must be presented as notation, not
finding.
Under performance pressure, effort shifts from improvement work to firefighting, and because eroded capability generates more problems after a delay, the shift is self-reinforcing: organizations settle into a low-capability, high-firefighting equilibrium that recurs across manufacturing, product development, facilities, and software.
supportsprimary-checked§Shortcuts and the Capability Trap (MIT-hosted PDF)
“The Reinvestment loop means a temporary emphasis on one option at the expense of the other is likely to be reinforced and eventually become permanent.”
Repenning, N., Sterman, J. (2001). Nobody Ever Gets Credit for Fixing Problems that Never Happened: Creating and Sustaining Process Improvement. California Management Review, 43(4), 64-88. doi:10.2307/41166101
supportsprimary-checkedDu Pont modeler interview quote (MIT-hosted PDF)
“As soon as you get the problems down, people will be taken away from the effort and the problems will go back up.”
Repenning, N., Sterman, J. (2001). Nobody Ever Gets Credit for Fixing Problems that Never Happened: Creating and Sustaining Process Improvement. California Management Review, 43(4), 64-88. doi:10.2307/41166101
“well-intentioned efforts by managers to search locally for the optimal workload balance lead them to systematically overload their organization and, thereby, cause capabilities to erode”
capability trajectories under local workload search: diverging trajectories; erosion rates model-derived (n = two software development organizations)
supportsprimary-checkedCase description (published PDF via MIT Sloan; both sentences re-verified verbatim 2026-07-09)
“Between 2005 and 2007, more than 85 percent of maintenance hours were spent responding to customer calls reporting problems or breakdowns. Approximately 40 percent were spent on urgent problems—those requiring a response within 2–3 days or sooner—compared to best practice benchmarks of 10 percent or less (Sullivan et al., 2010).”
reactive share of maintenance hours: >85% reactive share; the paper's own benchmark comparison is ~40% urgent vs best practice ≤10% (Sullivan et al., 2010) (n = one university campus, 2005-2007)
Lyneis, J., Sterman, J. (2016). How to Save a Leaky Ship: Capability Traps and the Failure of Win-Win Investments in Sustainability and Social Responsibility. Academy of Management Discoveries, 2(1), 7-32. doi:10.5465/amd.2015.0006
contextualizesprimary-checkedLead text (hbr.org)
“Serious problem-solving efforts degenerate into quick-and-dirty patching.”
Bohn, R. (2000). Stop Fighting Fires. Harvard Business Review, 78(4), July-August 2000. linknot peer-reviewed
supportsprimary-checkedpolicy tests / footnote 2
“The surge fails to lift the organization above the tipping point ($5 million/year increase in the R&M budget from 2010 to 2015); at least $15 million/year in additional R&M spending through 2030 is required”
Lyneis, J., Sterman, J. (2016). How to Save a Leaky Ship: Capability Traps and the Failure of Win-Win Investments in Sustainability and Social Responsibility. Academy of Management Discoveries, 2(1), 7-32. doi:10.5465/amd.2015.0006
Counter-evidence searched: Convergent formal models plus independent field studies across four industries, but small-N cases and simulations, some with author involvement, and no controlled trials — held at moderate per the seed report. The counter-search surfaced two bounds encoded as their own claims: the Nordhaus measurement-without-data critique of the underlying method (clm.firefighting-capability-trap.sd-models-caution) and lean's low-buffer success (clm.firefighting-capability-trap.lean-improvement-slack); both bound quantitative precision rather than contradicting the mechanism.
In formal models of multi-project development, firefighting has a tipping point set by steady-state utilization: a temporary workload increase of roughly 20% is absorbed and the system recovers, while roughly 25% pushes it past an unstable threshold into a permanent firefighting equilibrium that does not recover without intervention.
Repenning, N. (2001). Understanding Fire Fighting in New Product Development. Journal of Product Innovation Management, 18(5), 285-300. doi:10.1111/1540-5885.1850285
“And, as the experiment highlights, once the system reaches the lower equilibrium, absent an additional intervention, it never recovers.”
Repenning, N. (2001). Understanding Fire Fighting in New Product Development. Journal of Product Innovation Management, 18(5), 285-300. doi:10.1111/1540-5885.1850285
“in product development systems there exists a threshold for problem-solving activity that, when crossed, causes firefighting to spread rapidly from a few isolated projects to the entire development system”
Repenning, N., Gonçalves, P., Black, L. (2001). Past the Tipping Point: The Persistence of Firefighting in Product Development. California Management Review, 43(4). link
“Once the company is caught in such a "firefighting" mode, since it favors completing all current phase activities before undertaking any early phase work, it is stuck, continuing to operate in a lower-quality regime long after the workload returns to its original level.”
Black, L., Repenning, N. (2001). Why Firefighting Is Never Enough: Preserving High-Quality Product Development. System Dynamics Review, 17(1), 33-62. doi:10.1002/sdr.205
Counter-evidence searched: Held at suggestive despite four primary-checked excerpts because every number is a model illustration, not a measured constant — the authors say so themselves, and the Nordhaus-line critique of system-dynamics calibration applies. No field measurement of an organizational tipping threshold was found. Attribution note: the ~60%→70% (recovering) vs ~40%→25% (decaying) one-year map values that the seed report's table attributes to Black & Repenning were verified verbatim in Repenning's JPIM working paper; the verified source pack got them via a third-party reading (metasd) of Black & Repenning.
Escaping the capability trap requires a worse-before-better period in which shifting effort to improvement first raises costs and unresolved problems before capability recovers, and escape attempts fail when the investment is too small, too brief, or abandoned during the transitional dip.
supportsprimary-checkedBP Lima refinery case (MIT-hosted PDF); program start verified in the same passage: 'In 1994, the Lima facility introduced the maintenance learning lab and other system dynamics tools.'
“During the first six months, maintenance costs ballooned by 30%.”
transitional cost of escape at BP Lima: +30% maintenance costs in six months; pump MTBF then rose 12→58 months, pump failures >640 (1991) → 131 (1998) (n = one refinery)
Repenning, N., Sterman, J. (2001). Nobody Ever Gets Credit for Fixing Problems that Never Happened: Creating and Sustaining Process Improvement. California Management Review, 43(4), 64-88. doi:10.2307/41166101
supportsprimary-checkedTable 1, 'Improvement at the Lima Refinery' (MIT-hosted PDF, re-verified 2026-07-07) — the 640→131 series is PUMP failures specifically, and its 1991 baseline predates the 1994 program start by three years
“Pump MTBF up from 12 to 58 months (failures down from more than 640 in 1991 to 131 in 1998).”
pump failures at BP Lima: >640 (1991) → 131 (1998), ≈×4.8 reduction; window opens three years before the program began (n = one refinery)
Repenning, N., Sterman, J. (2001). Nobody Ever Gets Credit for Fixing Problems that Never Happened: Creating and Sustaining Process Improvement. California Management Review, 43(4), 64-88. doi:10.2307/41166101
supportsprimary-checkedModel analysis (published PDF via MIT Sloan)
“If a new effort begins (at time t2) and is not abandoned, the initial cost increase and performance drop eventually reverse, leading to lower costs and higher uptime, output, quality, reliability, and safety, in a worse-before-better pattern.”
Lyneis, J., Sterman, J. (2016). How to Save a Leaky Ship: Capability Traps and the Failure of Win-Win Investments in Sustainability and Social Responsibility. Academy of Management Discoveries, 2(1), 7-32. doi:10.5465/amd.2015.0006
supportsprimary-checkedDiscussion (published PDF via MIT Sloan); in the case, a +$5M/yr surge fails while ≥$15M/yr through 2030 crosses the threshold
“Escaping the trap requires investments large enough and sustained long enough to cross tipping thresholds that convert the vicious cycle into a virtuous cycle of better performance, greater investment, and still better performance.”
Lyneis, J., Sterman, J. (2016). How to Save a Leaky Ship: Capability Traps and the Failure of Win-Win Investments in Sustainability and Social Responsibility. Academy of Management Discoveries, 2(1), 7-32. doi:10.5465/amd.2015.0006
“In practice, however, obtaining additional resources can pose more problems than simply canceling projects or scaling down scope of work underway.”
Black, L., Repenning, N. (2001). Why Firefighting Is Never Enough: Preserving High-Quality Product Development. System Dynamics Review, 17(1), 33-62. doi:10.1002/sdr.205
Counter-evidence searched: The documented escapes are self-selected successes, partly narrated by an involved consultant (Ledet at Du Pont/BP Lima), with no failed-escape control sample; the MIT-facilities evidence pairs a panel regression with simulation in a single organization. Held at moderate because two independent, primary-checked case literatures plus the model mechanism converge; a survivorship-focused search found no systematic study of abandoned escapes. Scope notes (r1): the Lima 640→131 series is pump failures (Table 1), not refinery-wide equipment failures, and its 1991 baseline predates the 1994 program start; no cited source supports abandoned escapes settling permanently below their starting baseline.
Visible fixes earn systematically more credit than invisible prevention: voters reward disaster relief spending but give essentially no electoral reward to preparedness spending despite a roughly 15:1 value ratio, and action bias makes conspicuous intervention feel more defensible than optimal inaction.
supportsprimary-checkedAbstract (Cambridge Core)
“voters reward the incumbent presidential party for delivering disaster relief spending, but not for investing in disaster preparedness spending”
electoral reward for preparedness vs relief spending: relief rewarded; preparedness ~zero reward (n = U.S. county-level panel, 1988-2004)
Healy, A., Malhotra, N. (2009). Myopic Voters and Natural Disaster Policy. American Political Science Review, 103(3), 387-406. doi:10.1017/S0003055409990104
supportsprimary-checkedAbstract (Cambridge Core)
“We estimate that $1 spent on preparedness is worth about $15 in terms of the future damage it mitigates.”
value of preparedness spending: $1 preparedness ≈ $15 mitigated future damage
Healy, A., Malhotra, N. (2009). Myopic Voters and Natural Disaster Policy. American Political Science Review, 103(3), 387-406. doi:10.1017/S0003055409990104
“a goal scored yields worse feelings for the goalkeeper following inaction (staying in the center) than following action (jumping), leading to a bias for action”
goalkeeper behavior vs optimum: stay-center optimal (stop rate 33.3% vs 14.2% left / 12.6% right) yet chosen on ~6% of kicks (n = 286 penalty kicks; survey of 32 elite goalkeepers)
Bar-Eli, M., Azar, O., Ritov, I., Keidar-Levin, Y., Schein, G. (2007). Action bias among elite soccer goalkeepers: The case of penalty kicks. Journal of Economic Psychology, 28(5), 606-621. doi:10.1016/j.joep.2006.12.001
supportsreport-derivedAbstract, as quoted in the verified source pack (Springer paywalled)
“Individuals have a penchant for action, often for good reasons. But action bias arises if that penchant is carried over to areas where those reasons do not apply, hence is nonrational.”
Patt, A., Zeckhauser, R. (2000). Action Bias and Environmental Decisions. Journal of Risk and Uncertainty, 21(1), 45-72. doi:10.1023/A:1026517309871
“most organizations reward last-minute problem solving over the learning, training, and improvement activities that prevent such crises in the first place”
Repenning, N., Sterman, J. (2001). Nobody Ever Gets Credit for Fixing Problems that Never Happened: Creating and Sustaining Process Improvement. California Management Review, 43(4), 64-88. doi:10.2307/41166101
Counter-evidence searched: Capped at moderate per the verified source packs' honest gap statement: no controlled experiment was found in which evaluators rate or promote a 'fixer' above a matched 'preventer'. The asymmetry rests on one electoral panel, action-bias experiments, and field/case evidence, and the domains (elections, sports, policy vignettes) transfer to organizational promotion decisions only by analogy. Wording note: the seed report and source pack both render the $15 sentence as 'every dollar spent on preparedness is worth about $15 in terms of future damage mitigation'; the actual abstract wording (verified at Cambridge Core) is used here.
In Perlow's nine-month software-team ethnography, formal rewards tracked managers' perceptions of individual heroics in visible crises — top-ranked engineers were annotated for 80-100 hour weeks — while a $5,000 prevention request went unfunded for four months and a quiet-time intervention's gains faded because the reward structure was unchanged.
supportsprimary-checked§Rewards based on individual heroics (open PDF at interruptions.net)
“Ultimately, an engineer's rewards depended on managers' perceptions of the individual's heroics, as demonstrated by doing whatever it took to solve high-visibility crises at work; one's work process and its implications for others, whether unhelpful or disruptive, did not matter.”
Perlow, L. (1999). The Time Famine: Toward a Sociology of Work Time. Administrative Science Quarterly, 44(1), 57-81. doi:10.2307/2667031
supportsprimary-checkedYear-end ranking evidence (open PDF at interruptions.net)
“The comment following the engineer ranked first on the list read: "works 80-100 hours/week."”
Perlow, L. (1999). The Time Famine: Toward a Sociology of Work Time. Administrative Science Quarterly, 44(1), 57-81. doi:10.2307/2667031
supportsprimary-checkedHelp-line vignette (open PDF at interruptions.net)
“For four months, as she continued to slip further and further behind schedule, the engineer repeatedly asked for the purchase of the help line. Only in December, when her situation became desperate, did the relevant decision makers at Ditto agree to spend the $5,000.”
Perlow, L. (1999). The Time Famine: Toward a Sociology of Work Time. Administrative Science Quarterly, 44(1), 57-81. doi:10.2307/2667031
“Our [company] culture rewards the heroes. Frankly, that's how I got where I've gotten.”
Repenning, N., Sterman, J. (2001). Nobody Ever Gets Credit for Fixing Problems that Never Happened: Creating and Sustaining Process Improvement. California Management Review, 43(4), 64-88. doi:10.2307/41166101
contextualizesprimary-checkedData Sources
“I studied the software group over the product's nine-month development cycle, from the commitment of funding until the product's launch.”
Perlow, L. (1999). The Time Famine: Toward a Sociology of Work Time. Administrative Science Quarterly, 44(1), 57-81. doi:10.2307/2667031
Managers tend to attribute capability-trap symptoms to worker effort or ability rather than to system structure, and because cutting improvement genuinely raises short-run output, the resulting get-tough policies appear to work and confirm the misattribution.
“the critical determinants of success in efforts to learn and improve are the interactions between managers' attributions regarding the cause of poor organizational performance and the physical structure of the workplace”
Repenning, N., Sterman, J. (2002). Capability Traps and Self-Confirming Attribution Errors in the Dynamics of Process Improvement. Administrative Science Quarterly, 47(2), 265-295. doi:10.2307/3094806
supportsprimary-checkedWorked example (MIT-hosted PDF)
“Because managers do not fully observe the reduction in training, experimentation, and improvement effort (they fail to account for the Shortcuts loop), they overestimate the impact of their get-tough policy, in our example by as much as a factor of three.”
Repenning, N., Sterman, J. (2001). Nobody Ever Gets Credit for Fixing Problems that Never Happened: Creating and Sustaining Process Improvement. California Management Review, 43(4), 64-88. doi:10.2307/41166101
contextualizesprimary-checkedAbstract (via ERIC record EJ318293); ASQ full text paywalled
“The attributional perspective on leadership, which suggests the social construction of organizational realities attributes to leadership the activities and outcomes of organizations, was supported by the results of three archival studies and a series of experimental studies.”
Meindl, J., Ehrlich, S., Dukerich, J. (1985). The Romance of Leadership. Administrative Science Quarterly, 30(1), 78-102. doi:10.2307/2392813
supportsprimary-checkedAbstract, opening sentence (publisher-deposited record, read via OpenAlex W2001253236 and the CoLab mirror; the deposit itself truncates mid-word, and the full abstract is paywalled)
“Suggesting that as an explanatory concept, leadership has assumed a heroic, larger-than-life quality, this research investigated the effects of leadership attributions on evaluations of organizatio…”
Meindl, J., Ehrlich, S. (1987). The Romance of Leadership and the Evaluation of Organizational Performance. Academy of Management Journal, 30(1), 91-109. doi:10.2307/255897
supportsreport-derivedAs summarized in Bligh & Kohles (2009), 'Romance of Leadership', Encyclopedia of Group Processes & Intergroup Relations, pp. 718-720 (open SAGE reference PDF). The entry does not cite the 1987 paper by name (its reference list carries only the 1985 ASQ article); identifying the 'subsequent research' as this 1987 follow-up (AMJ 30(1), 91-109) is our inference, consistent with the 1987 abstract's opening and with Hammond et al. (2021) naming this experiment as their replication target.
“subsequent research has demonstrated that people value performance results more highly when those results are attributed to leadership and that a halo effect exists for leadership”
Meindl, J., Ehrlich, S. (1987). The Romance of Leadership and the Evaluation of Organizational Performance. Academy of Management Journal, 30(1), 91-109. doi:10.2307/255897
contradictsprimary-checkedAbstract (via SJSU ScholarWorks record); four studies, of which Studies 1 and 2 are close replications
“do not support Meindl and Ehrlich's findings that organizations are viewed more favorably when such outcomes are attributed to leadership”
replication of the 1987 evaluation experiment: null result across four studies (two close, two conceptual replications)
Hammond, M., Schyns, B., Lester, G., Clapp-Smith, R., Thomas, J. (2021). The Romance of Leadership: Rekindling the fire through replication of Meindl and Ehrlich. The Leadership Quarterly, 32(6), 101538. doi:10.1016/j.leaqua.2021.101538
Counter-evidence searched: The attribution error is modeled and consistent with case interviews but has not been measured psychometrically in trap settings (the seed report's own rating: theoretical/suggestive). Round-1 correction (2026-07-07): the identical-outcomes-evaluated-more-favorably experimental finding had been mis-attributed to the 1985 ASQ paper; it belongs to the Meindl & Ehrlich 1987 AMJ follow-up, now its own source. The counter-search on that arm found Hammond et al. (2021): four replication studies that do not support the 1987 evaluation effect, so the romance-of-leadership lab arm is carried as weakened and the 1985 archival attribution evidence as the surviving pillar. No dedicated search yet for evidence that managers correctly attribute trap symptoms.
In practitioner doctrine, rising dependence on heroic out-of-process intervention is treated as a leading indicator of systemic overload: Google SRE holds that heroism suppresses the broken-system signal, and in its published team case heroic paging response preceded the formal declaration of operational overload.
“A signal that a system is broken is a really valuable signal.”
Malmberg, A. (2024). Why heroism is bad, and what we can do to stop it. Google SRE resources (sre.google). linknot peer-reviewed
supportsprimary-checkedWorkbook ch. 8 'On-Call', Connection-team case (~5 incidents/shift vs a budget of 2; engineers quit before overload was declared)
“Members of the team heroically responded to the daily onslaught of pages but couldn't keep up; there simply was not enough time in the day to find the root cause and properly fix the incoming issues.”
Beyer, B., Murphy, N., Rensin, D., Kawahara, K., Thorne, S. (2018). The Site Reliability Workbook: Practical Ways to Implement SRE. O'Reilly Media. linknot peer-reviewed
contextualizesprimary-checkedAbstract (open-access PMC copy) — the peer-reviewed cultural mechanism: how a culture processes bad news predicts its trouble handling
“Because information flow is both influential and also indicative of other aspects of culture, it can be used to predict how organisations or parts of them will behave when signs of trouble arise.”
Westrum, R. (2004). A typology of organisational cultures. Quality and Safety in Health Care, 13(Suppl II), ii22-ii27. doi:10.1136/qshc.2003.009522
supportsprimary-checkedSRE Workbook, ch. 8 (Connection team case)
“They had an established pager budget of two paging incidents per shift, but for the past year they had regularly been receiving five paging incidents per shift”
Beyer, B., Murphy, N., Rensin, D., Kawahara, K., Thorne, S. (2018). The Site Reliability Workbook: Practical Ways to Implement SRE. O'Reilly Media. linknot peer-reviewed
Counter-evidence searched: Single company, unquantified, stated as doctrine rather than tested — kept at anecdotal even though every excerpt is primary-checked. The HRO literature's bound is encoded separately (clm.firefighting-capability-trap.hro-response-designed-virtue): skilled response per se is not pathological. The strong predictive form — hero-celebration as a measurable distance-to-collapse metric — has never been tested longitudinally; no study linking reward-system composition to subsequent failure rates was found. Wording note: the seed report and source pack render the first quote without its leading 'Because'.
In DORA's surveys, high-performing software organizations report more time on new work and less on unplanned work and rework than low performers (49%/21% vs 38%/27%), and elite performers show large delivery-outcome multiples — but the gradient is not monotonic, because medium performers report more rework than low performers.
supportsprimary-checkedp. 26
“High performers reported spending 49 percent of their time on new work and 21 percent on unplanned work or rework. By contrast, low performers spend 38 percent of their time on new work and 27 percent on unplanned work or rework.”
time allocation, high vs low performers: new work 49% vs 38%; unplanned/rework 21% vs 27% (n = survey program spanning 25,000+ respondents over five years)
Brown, A., Forsgren, N., Humble, J., Kersten, N., Kim, G. (2016). 2016 State of DevOps Report. Puppet + DORA. linknot peer-reviewed
supportsprimary-checkedSDO performance section
“The mean between these two ranges shows a 7.5 percent change failure rate for elite performers and 53 percent for low performers.”
elite vs low delivery outcomes (2018): 46x deploy frequency; 2,555x faster lead time; 2,604x faster restore; 7x lower change failure rate (7.5% vs 53%) (n = ~1,900 respondents (2018); 30,000+ across five years)
Forsgren, N., Humble, J., Kim, G. (2018). Accelerate: State of DevOps 2018 - Strategies for a New Economy. DORA / Google Cloud. linknot peer-reviewed
“When they happen multiple times a year, they can quickly take over the work of the team so that unplanned work becomes the norm, leading to burnout, an important consideration for teams and leaders.”
Forsgren, N., Humble, J., Kim, G. (2018). Accelerate: State of DevOps 2018 - Strategies for a New Economy. DORA / Google Cloud. linknot peer-reviewed
contradictsprimary-checkedp. 41 — note the construct wording: p. 41 says 'rework' where p. 26 says 'unplanned work or rework' for the same 27% low-performer figure; the report does not explain the difference
“By contrast, we found that low performers spend less time on rework than medium performers (27 percent vs. 32 percent, respectively) and more time on new work (38 percent vs. 34 percent, respectively).”
rework and new work, medium vs low performers: rework 32% vs 27%; new work 34% vs 38% — the reactive-work gradient is non-monotonic at the bottom
Brown, A., Forsgren, N., Humble, J., Kersten, N., Kim, G. (2016). 2016 State of DevOps Report. Puppet + DORA. linknot peer-reviewed
Counter-evidence searched: Self-report, cross-sectional, vendor-sponsored surveys with constructed performance clusters — strong on breadth, weak on causal identification. The medium-vs-low rework anomaly is encoded as contradicting evidence; DORA's own explanation (low performers ignoring critical rework and racking up technical debt) is plausible but untested. Stat-vintage correction from the verified source pack confirmed against both PDFs: 200x/24x/3x/2,555x are 2016 headline multiples, not 2018; the 2018 elite-vs-low multiples are 46x/2,555x/2,604x/7x.
Sustained margin erosion ends in normalization of deviance: at NASA, repeated success with O-ring and foam anomalies reclassified them as acceptable, and the Columbia investigation concluded that organizational culture and structure had as much to do with the accident as the physical cause.
supportsprimary-checkedUniversity of Chicago Press book page
“NASA insiders, when repeatedly faced with evidence that something was wrong, normalized the deviance so that it became acceptable to them”
Vaughan, D. (1996). The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press. linknot peer-reviewed
supportsprimary-checkedVol. I, ch. 7, p. 177
“In the Board's view, NASA's organizational culture and structure had as much to do with this accident as the External Tank foam.”
Columbia Accident Investigation Board (2003). Columbia Accident Investigation Board Report, Volume I. NASA / U.S. Government Printing Office. linknot peer-reviewed
supportsprimary-checkedVol. I, ch. 6, p. 130 (passage citing Vaughan by name)
“The history of foam-problem decisions shows how NASA first began and then continued flying with foam losses, so that flying with these deviations from design specifications was viewed as normal and acceptable.”
Columbia Accident Investigation Board (2003). Columbia Accident Investigation Board Report, Volume I. NASA / U.S. Government Printing Office. linknot peer-reviewed
supportsprimary-checkedOpen-access PMC copy
“What begin as deviations from standard operating rules become, with enough repetitions, 'normalized' practice patterns.”
Banja, J. (2010). The normalization of deviance in healthcare delivery. Business Horizons, 53(2), 139-148. doi:10.1016/j.bushor.2009.10.006
Organizational slack has an inverse U-shaped relationship with innovation — both too little and too much slack are detrimental — so the capability-trap prescription must protect improvement allocation specifically rather than organizational fat in general.
“There is an inverse U-shaped relationship between slack and innovation in organizations: both too much and too little slack may be detrimental to innovation.”
slack ↔ innovation: inverse U-shaped; peak at intermediate slack (n = 264 functional departments, two multinational corporations)
Nohria, N., Gulati, R. (1996). Is Slack Good or Bad for Innovation?. Academy of Management Journal, 39(5), 1245-1264. doi:10.5465/256998
Deliberately low-buffer lean production systems outperformed buffered mass producers — NUMMI ran near-Toyota productivity and quality on two days of parts inventory — because they removed buffer slack while institutionalizing improvement slack (kaizen teams, andon stop authority), so low slack per se is not the capability-trap disease; reactive allocation of attention is.
supportsprimary-checkedDiscussion of Krafcik 1986 comparisons, Exhibit 5 (author-hosted PDF)
“labor productivity, both corrected and uncorrected for differences in product and technology, was much higher at NUMMI than at the old GM-Fremont plant in 1978 and at the GM-Framingham plant”
assembly hours per unit (Krafcik 1986): NUMMI ~20.8 vs GM-Framingham ~40.7 (Toyota Takaoka ~18.0) (n = plant-level benchmarking)
Adler, P. (1993). The 'Learning Bureaucracy': New United Motor Manufacturing, Inc.. Research in Organizational Behavior, vol. 15 (eds. Staw & Cummings), 111-194. linknot peer-reviewed
supportsprimary-checkedCase narrative (author-hosted PDF); the same passage notes this was still above Takaoka's two-hour level
“NUMMI parts inventories averaged two days.”
Adler, P. (1993). The 'Learning Bureaucracy': New United Motor Manufacturing, Inc.. Research in Organizational Behavior, vol. 15 (eds. Staw & Cummings), 111-194. linknot peer-reviewed
supportsreport-derivedSeed report, Key sources table (book not fetched)
“Lean (deliberately low-buffer) plants dramatically outperformed buffered mass producers on productivity and quality”
Womack, J., Jones, D., Roos, D. (1990). The Machine That Changed the World. Rawson Associates / Macmillan. linknot peer-reviewed
“Our sample-wide results indicate that JIT adopters improve financial performance relative to non-adopters, and that profit margin, rather than asset turnover, is the primary source of such improvement.”
Kinney, M., Wempe, W. (2002). Further Evidence on the Extent and Origins of JIT's Profitability Effects. The Accounting Review, 77(1), 203-225. doi:10.2308/accr.2002.77.1.203
“results of additional analyses suggest that JIT adopters below a firm-size threshold do not improve financial performance”
Kinney, M., Wempe, W. (2002). Further Evidence on the Extent and Origins of JIT's Profitability Effects. The Accounting Review, 77(1), 203-225. doi:10.2308/accr.2002.77.1.203
contextualizesprimary-checkedsentence preceding the two-days NUMMI inventory comparison
“the GM facilities, including Fremont, were all designed to stock several weeks of parts”
Adler, P. (1993). The 'Learning Bureaucracy': New United Motor Manufacturing, Inc.. Research in Organizational Behavior, vol. 15 (eds. Staw & Cummings), 111-194. linknot peer-reviewed
Counter-evidence searched: Counter-evidence to naive slack-hoarding readings of the trap thesis — and simultaneously a confirmation of its core, since lean pairs low buffers with obsessive institutionalized improvement. Bounds: benchmarking and observational adoption studies, not causal identification; Kinney & Wempe's firm-size threshold shows the payoff is conditional; the buffer-slack vs improvement-slack distinction follows the verified source pack's synthesis of Adler's NUMMI chapter.
High-reliability organizations treat skilled crisis response (resilience and containment) as a designed capability alongside anticipation because not all surprises are preventable, so the presence of skilled emergency response is not itself evidence of a capability trap — reward-dominant, chronic dependence on it is.
“Five principles: anticipation (preoccupation with failure, reluctance to simplify, sensitivity to operations) + containment (commitment to resilience, deference to expertise). Resilience = capacity to "absorb strain and preserve function despite adversity."”
Weick, K., Sutcliffe, K. (2015). Managing the Unexpected: Sustained Performance in a Complex World (3rd ed.). Wiley. doi:10.1002/9781119175834not peer-reviewed
contextualizesprimary-checkedsre.google 'no-heroes' essay — the practitioner line between designed response capacity and chronic heroics
“If you're committing to an SLO, you should be able to meet the SLO without any heroes.”
Malmberg, A. (2024). Why heroism is bad, and what we can do to stop it. Google SRE resources (sre.google). linknot peer-reviewed
The system-dynamics tradition behind the capability-trap models has faced sustained methodological criticism — Nordhaus's review of Forrester's World Dynamics found parameters asserted without empirical estimation — so simulated trap thresholds should be presented as hypothesis-generating illustrations rather than measured constants.
“Critical audit of Forrester's SD model: parameters not empirically estimated; conclusions driven by assumed feedback structures.”
Nordhaus, W. (1973). World Dynamics: Measurement Without Data. The Economic Journal, 83(332), 1156-1183. doi:10.2307/2230846
contextualizesprimary-checkedMIT working-paper version — the authors' own framing of the model as stylized
“So, to understand firefighting, consider the following stylized model of a product development system (Figure 1).”
Repenning, N., Gonçalves, P., Black, L. (2001). Past the Tipping Point: The Persistence of Firefighting in Product Development. California Management Review, 43(4). link
Counter-evidence searched: Forrester and colleagues published a reply (Policy Sciences, 1974) defending the approach, and the trap papers are considerably more field-grounded than the world models Nordhaus attacked — Lyneis & Sterman pair simulation with panel regression. The caution bounds quantitative precision (thresholds, rates), not the existence of the mechanism.
Google SRE institutionalizes countermeasures to the credit asymmetry by policy rather than exhortation: a 50% cap on operational toil, error budgets that turn the reliability/velocity trade-off into an explicit depoliticized quantity, and pager budgets that define overload before attrition does.
“At least 50% of each SRE's time should be spent on engineering project work that will either reduce future toil or add service features.”
Beyer, B., Jones, C., Petoff, J., Murphy, N. (2016). Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media. linknot peer-reviewed
“The error budget provides a clear, objective metric that determines how unreliable the service is allowed to be within a single quarter. This metric removes the politics from negotiations between the SREs and the product developers when deciding how much risk to allow.”
Beyer, B., Jones, C., Petoff, J., Murphy, N. (2016). Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media. linknot peer-reviewed
“Toil tends to expand if left unchecked and can quickly fill 100% of everyone's time.”
Beyer, B., Jones, C., Petoff, J., Murphy, N. (2016). Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media. linknot peer-reviewed
supportsprimary-checkedWorkbook ch. 8 'On-Call'
“We target a maximum of two incidents per on-call shift, to ensure adequate time for follow-up.”
Beyer, B., Murphy, N., Rensin, D., Kawahara, K., Thorne, S. (2018). The Site Reliability Workbook: Practical Ways to Implement SRE. O'Reilly Media. linknot peer-reviewed