Packet Path
Trace endpoint access, representative fabric lanes and one shared meet-me-room boundary without pretending to size a real network.
What this experiment teaches
- Separate workload fabric from external reach
- Expose a shared carrier boundary
- See access-placement implications without a hidden winner
This standalone teaching model isolates one mechanism so its inputs and consequence remain inspectable. The illustrative label describes the output kind; it does not turn the result into measured facility data, product selection, a commissioned design or professional advice.
Predict before changing the model
Predict whether the selected failure removes local workload connectivity, external reach, or both.
Record the expected direction in your own words. Then change one input at a time and compare the result with that prediction. If the output surprises you, open the model card and inspect the boundary and assumptions before creating a story around the number or state.
Stress and debrief
Change access placement, preserve the same rack count and compare the modeled switch and horizontal-run implications.
A useful debrief names four things separately: the input that changed, the modeled consequence, the important effects that were not evaluated and the site data or engineering work needed before a real decision. The lab deliberately avoids a universal score because energy, resilience, capacity, maintainability and risk are different dimensions.
An SLA is a measured commitment
A service-level agreement defines the service, measurement window, exclusions and consequence of missing the target. “99.9%” alone is incomplete: the endpoint, clock, calculation, maintenance rules and remedy all matter.
Availability percentages become small downtime budgets at high targets. Operators therefore need monitoring that measures the same service the customer experiences, not only green infrastructure lights.
Recovery time is an engineering variable
MTBF summarizes the average operating interval between repairable failures. MTTR usually describes mean time to repair or restore, but teams must define which. Faster detection, clear runbooks, accessible spares, practiced escalation and reversible change can reduce recovery time.
A high-reliability component can still produce a long outage if diagnosis or replacement is slow. Conversely, a component that fails more often may have limited service impact when it is isolated and quickly restored.
Fault domains contain blast radius
A fault domain is the set of resources expected to fail together: a rack, power path, network zone, cooling loop, software cluster or site. Redundancy should cross the boundary of the failure it is meant to survive.
Document dependencies and test assumptions. Two services in different racks may still share a switch, PDU, control plane or change process. Common-mode and human failures frequently ignore labels drawn on a diagram.
Continue the evidence trail
Open the connected Academy lesson. The lesson provides the formula or mechanism, common misconception, knowledge check and source context that surround this compact experiment.
- Diverse carrier entrances — Two carriers are not physically diverse when their routes converge in the same trench, building entrance, riser or room.
- Meet-me room boundary — The meet-me room organizes carrier handoff and cross-connects, but a single room can become a shared physical fault domain.
- Leaf-spine fabric — A leaf-spine fabric creates repeatable east-west paths when every leaf reaches every spine and the routing design uses those paths.
- Three-tier network — Access, aggregation and core layers create distinct network roles, but traffic paths and oversubscription depend on the actual design.
Starting references
- NIST SP 800-34 Rev. 1 — Contingency Planning Guide
- Uptime Institute — Data Center Tier Classification
Current standards, adopted requirements, verified site information, manufacturer data and qualified professional review remain necessary for real work.