Skip to content
LAB-06 · RESILIENCE · CALCULATED MODEL

Availability Budget

See how a small percentage difference becomes a sharply smaller annual unavailability allowance.

What this experiment teaches

  • Translate percentage into minutes
  • Keep the reporting period explicit
  • Avoid treating an SLA target as a topology result

This standalone teaching model isolates one mechanism so its inputs and consequence remain inspectable. The calculated label describes the output kind; it does not turn the result into measured facility data, product selection, a commissioned design or professional advice.

Predict before changing the model

Predict the order-of-magnitude change in downtime when another nine is added to the target.

Record the expected direction in your own words. Then change one input at a time and compare the result with that prediction. If the output surprises you, open the model card and inspect the boundary and assumptions before creating a story around the number or state.

Stress and debrief

Compare the same total downtime as one long outage versus many short incidents and list what this simple budget cannot distinguish.

A useful debrief names four things separately: the input that changed, the modeled consequence, the important effects that were not evaluated and the site data or engineering work needed before a real decision. The lab deliberately avoids a universal score because energy, resilience, capacity, maintainability and risk are different dimensions.

An SLA is a measured commitment

A service-level agreement defines the service, measurement window, exclusions and consequence of missing the target. “99.9%” alone is incomplete: the endpoint, clock, calculation, maintenance rules and remedy all matter.

Availability percentages become small downtime budgets at high targets. Operators therefore need monitoring that measures the same service the customer experiences, not only green infrastructure lights.

Recovery time is an engineering variable

MTBF summarizes the average operating interval between repairable failures. MTTR usually describes mean time to repair or restore, but teams must define which. Faster detection, clear runbooks, accessible spares, practiced escalation and reversible change can reduce recovery time.

A high-reliability component can still produce a long outage if diagnosis or replacement is slow. Conversely, a component that fails more often may have limited service impact when it is isolated and quickly restored.

Fault domains contain blast radius

A fault domain is the set of resources expected to fail together: a rack, power path, network zone, cooling loop, software cluster or site. Redundancy should cross the boundary of the failure it is meant to survive.

Document dependencies and test assumptions. Two services in different racks may still share a switch, PDU, control plane or change process. Common-mode and human failures frequently ignore labels drawn on a diagram.

Continue the evidence trail

Open the connected Academy lesson. The lesson provides the formula or mechanism, common misconception, knowledge check and source context that surround this compact experiment.

Starting references

Current standards, adopted requirements, verified site information, manufacturer data and qualified professional review remain necessary for real work.