• Docs
  • Talk to an expert
Blog
Blog
BlogProductCase studiesNewsInsights
Blog

That 2am incident cost you more than downtime

preview environmentsdeveloper workflowDevOpsdata cloning
07 October 2026
Share

TL;DR

  • The downtime number is what shows up on the status page. It's not what the incident actually cost you.
  • Whoever gets pulled in doesn't just lose the incident time, they lose whatever they were building before the pager went off, and that never makes it into the postmortem.
  • Fix the reproduction step and most of this cost goes away on its own.

It's 2:14am. A pager alert wakes an on-call engineer for a checkout failure hitting a slice of customers. By the time they've pulled the right logs, cross-checked the deploy history, and finally gotten the bug to happen again in front of them, the sun's coming up. 

The incident report will list two hours of downtime. It won't say anything about the day that engineer just lost, or the fact that nobody on your team could have told you in advance how long that reproduction step was going to take.

The bill you don't see on the status page

Key takeaway: downtime is the cost customers see, reproduction time is the cost your team quietly absorbs.

Time-to-resolution gets tracked on every incident dashboard. Time-to-reproduce almost never does, and it's usually the bigger of the two. Before anyone can fix anything, someone has to make the bug happen again on purpose, with roughly the right data and roughly the right traffic hitting roughly the right combination of service versions that only exists in production. None of that shows up on a dashboard. It's also most of where the hours go.

Reproduction is the bottleneck, not diagnosis

Key takeaway: most of the time goes into getting the bug to show up, not fixing it once it does.

Once something reproduces reliably, the fix is usually fast. What eats the time is everything before that: hunting for a data snapshot, staring at a diff between staging and prod trying to spot what changed, or just waiting around for the bug to happen again on its own. Local environments are supposed to mirror production. In practice they rarely do, and that gap is where the hours quietly disappear. 

It also shows up in the numbers: DORA's 2024 State of DevOps report puts elite teams at under an hour to recover from a failed deployment, while low performers take anywhere from a week to a month. Reproduction speed is most of what separates those two numbers, and the gap isn't closing. DORA's data shows the share of low-performing teams grew from 17% to 25% between 2023 and 2024, while the high-performing tier shrank from 31% to 22% over the same period.

What context switching actually costs the roadmap

Key takeaway: the interruption costs twice, once when the engineer drops their task and again when they have to rebuild the context to pick it back up.

Getting pulled into an incident doesn't necessarily cost time the way you'd expect. In a study of interrupted work, Gloria Mark and colleagues at UC Irvine and Humboldt University found that people who were interrupted mid-task actually finished faster than people working uninterrupted, not because the interruption was free, but because people compensate by working faster and writing less. 

That compensation had a real cost: the interrupted group reported significantly higher stress, frustration, time pressure, and effort than the uninterrupted group. None of that shows up in a completion-time metric, and none of it makes it into the postmortem either.

Why "works on my machine" persists

Key takeaway: giving experienced teams a safe way to reproduce production removes the guesswork that's actually slowing them down.

Experienced teams hit this wall too. Their tooling doesn't give them a safe, disposable copy of production to test against, so the fallback is reading logs and hoping the bug shows up again. 

A preview environment, a full clone of production spun up on demand, is what actually removes the guesswork, full stop.

See how a full-stack preview environment spins up in seconds. 


 

Frequently asked questions (FAQ)

Why does incident cost look bigger after the fact than during it? 
The downtime clock stops fast. The reproduction and context-switch clock keeps running long after the status page says "resolved."

Is this a tooling problem or a process problem? 
Mostly tooling. A good process can only work around a slow reproduction step, it can't remove it.

What actually shortens time-to-reproduce? 
A preview environment that matches production's data and configuration on demand, instead of guessing locally or waiting for the bug to happen again.

Stay updated

Subscribe to our monthly newsletter for the latest updates and news.

Deploy possibility.
Try Upsun for free.

Build with DispatchDeploy on Cloud