Skip to content
OPERATOR ACADEMY · GLOSSARY · DEFINED TERM

MTTR

Mean time to repair or restore: the average time from a failure to the service being usable again.

Also called: Mean time to repair, Mean time to restore

In practice

MTTR is the half of availability an organisation can most directly improve. It covers everything between a failure occurring and the service working again — detection, notification, diagnosis, getting the right person and the right part to the right place, the repair itself, and verifying that the fix worked.

Measured honestly, the repair is rarely the largest part. Time spent not knowing there is a problem, time spent working out which of several symptoms is the cause, and time spent waiting for access, a part or an authorisation usually dominate. That is good news, because those are addressed with monitoring, runbooks, spares strategy and clear on-call arrangements rather than with capital.

The distinction between repair and restore is worth keeping. Restoring service by failing over to another path can happen in seconds; repairing the failed component may take days. An availability figure cares about the first; a maintenance plan cares about the second. Quoting one while meaning the other is a common source of disagreement.

Because MTTR sits in the availability relationship alongside MTBF, halving it has a comparable effect to doubling the interval between failures — and it is usually the cheaper of the two to achieve.

Scope of this definition

This is an educational summary of how the term is used in data center practice. It is not a standard, a specification or engineering advice, and where a real decision depends on it, the current adopted standards, verified site information and qualified professional review are the correct sources.

Back to the A–Z glossary · Operator Academy

Related material

Related terms

  • MTBF — Mean time between failures: the average interval between failures of a repairable item.
  • RTO — Recovery time objective: the target duration within which a service should be restored after a disruption.
  • Fault domain — A set of resources expected to be affected by the same underlying failure.

Where this term is used

Reference