MTTR
Mean time to repair or restore: the average time from a failure to the service being usable again.
Also called: Mean time to repair, Mean time to restore
In practice
MTTR is the half of availability an organisation can most directly improve. It covers everything between a failure occurring and the service working again — detection, notification, diagnosis, getting the right person and the right part to the right place, the repair itself, and verifying that the fix worked.
Measured honestly, the repair is rarely the largest part. Time spent not knowing there is a problem, time spent working out which of several symptoms is the cause, and time spent waiting for access, a part or an authorisation usually dominate. That is good news, because those are addressed with monitoring, runbooks, spares strategy and clear on-call arrangements rather than with capital.
The distinction between repair and restore is worth keeping. Restoring service by failing over to another path can happen in seconds; repairing the failed component may take days. An availability figure cares about the first; a maintenance plan cares about the second. Quoting one while meaning the other is a common source of disagreement.
Because MTTR sits in the availability relationship alongside MTBF, halving it has a comparable effect to doubling the interval between failures — and it is usually the cheaper of the two to achieve.
Scope of this definition
This is an educational summary of how the term is used in data center practice. It is not a standard, a specification or engineering advice, and where a real decision depends on it, the current adopted standards, verified site information and qualified professional review are the correct sources.