Key Takeaways
- High availability can mask a lack of resilience when failures occur outside the architecture's design assumptions.
- Multi-region architectures can still share control-plane dependencies such as DNS health checks, identity and access management, and routing infrastructure.
- Resilience needs explicit ownership, recurring testing, and continuously maintained recovery procedures rather than relying on on-call ownership alone.
- The cost and operational risk of full failover testing can push organizations toward performative resilience, where recovery is assumed rather than demonstrated.
- Resilience in complex distributed systems is probabilistic, so the practical goal is to build confidence by repeatedly testing recovery paths and exposing hidden dependencies.
A team I worked with upgraded their public ingress load balancers to Transport Layer Security (TLS) 1.3 for compliance. Nothing about the rollout indicated a problem: handshakes completed, services remained healthy, and metrics remained flat.
As of this writing, Route 53 HTTPS health checks require the target endpoint to support TLS 1.2. If TLS 1.2 is disabled and only TLS 1.3 is available, the health checker cannot complete the handshake and will mark the endpoint unhealthy. The CDN deemed the region unhealthy and stopped routing traffic to it.
Within the region, everything looked fine. Services were up, load balancers were healthy, applications continued to serve requests, and dashboards showed nothing unusual. The only visible symptom was that traffic had stopped arriving.
Users were transparently rerouted to another region thousands of kilometers away. Latency spiked in affected geographies. The failover region started scaling under load for which it had not been provisioned. Synthetic monitoring eventually revealed a pattern that internal telemetry could not show. The actual issue took roughly forty minutes to isolate. The failure wasn't happening in the application data plane. It was happening in the control plane, which decides where traffic should go.
The system was highly available, but it was not resilient.
This distinction matters more than most cloud architectures acknowledge. In practice, the two terms are often used interchangeably, even in teams with a strong engineering culture. Yet they describe different problems with different prerequisites. High availability is about surviving expected failures with minimal interruption. Resilience, on the other hand, is about recovering from conditions the system was never explicitly designed to handle. Part of the confusion is that availability is measurable in terms of uptime percentages, failover timing, and replication lag. Resilience resists that kind of measurement. Many failure modes emerge only under real pressure. Meaningful recovery tests are often too expensive or too risky to deliberately reproduce.
Availability and Resilience Are Not the Same Problem
Modern cloud platforms make high availability relatively accessible. A managed database with Multiple Availability Zones (Multi-AZ) failover, autoscaling groups behind a load balancer, and a CDN in front helps most systems achieve respectable uptime. Standard patterns work well when failures stay within the assumptions the architecture was built around: an instance crashing, an availability zone disappearing, or a dependency timing out.
The problem is that real incidents rarely fit those assumptions. High-availability engineering assumes failures are isolated and predictable. Resilience engineering starts with the opposite belief. Eventually something critical will fail in a way nobody modelled. Redundancy alone will not protect against failure. A failover path that has never been exercised is not a recovery strategy; it is just an assumption. Three patterns make the gap concrete.
Correlated Failures Across Redundancy
Modern HA does account for some correlated failures, such as Multi-AZ, which clearly models an entire zone going down at once. What it doesn't account for is software-layer correlation, such as a bad config push deployed to all replicas simultaneously, a poisoned cache record served from every read replica, or a dependency upgrade that silently breaks a contract on which the whole system relies. When independence breaks down at that level, redundancy stops working as insurance.
Graceful Degradation That Isn’t
Most systems are assumed to degrade gracefully, but few have ever been tested under realistic loads. When the read replica falls behind under stress, does the cache layer absorb the pressure or amplify it onto the primary? Until exercised, graceful degradation is just a hypothesis, rather than a guarantee.
Recovery Paths That Have Rotted
The system has been running for two years; the rebuild runbook was written before half of the current dependencies existed. IAM policies have drifted and deployment tooling has changed. Recovery is a code path like any other; unused code paths rot.
The Organizational Problems Appear Long Before the Incident
The resilience problem is both architectural and organizational. The architectural part gets more attention. Diagrams are easier to draw than operational ownership is to execute.
I observed one recurring problem. Recovery ownership is often ambiguous even when uptime ownership is not. Most organizations are clear about who owns availability, such as on-call rotations, SLOs, and escalation channels. Few have equivalent structures for resilience. There is no named recovery owner, no recurring testing cadence, and no process for keeping runbooks current as infrastructure evolves. The gap is rarely a deliberate decision; it is the default when nobody claims ownership.
Part of what keeps that gap open is that closing it properly is expensive. Testing the scenarios that actually matter requires dedicated engineering time, operational risk, and infrastructure maintained solely for events that may never occur. The result is a drift toward performative resilience. The architecture diagram looks right, the standby region exists, and the runbook exists. Actual recovery capability is assumed rather than demonstrated, not because engineers are careless, but because proving it rigorously costs real money, carries real risk, and delivers value only at the most adverse moment.
In practice, the first thirty minutes of a major incident are often lost to dependency archaeology. Staff finds that permissions have drifted, runbooks reference tooling retired a year ago, and nobody is sure who still understands the standby environment. None of these problems is exotic. They are normal operational erosion. Recovery paths decay because they are rarely exercised.
The fix is unglamorous and requires explicit recovery ownership, which is separate from on-call. On-call is about response. Recovery ownership is about preparation, which includes the procedure, the tooling, and the responsibility to keep both current. Conflate the two and you get great incident commanders and runbooks that haven't been opened in eighteen months.
Teams that recover well treat failover drills as a recurring engineering commitment. They start small with one service and one AZ shift. Every exercise finds something broken, such as a revoked permission, a stale runbook, or a silent alarm. You find these issues by running the process. The shift is treating recovery as something you maintain, rather than something the architecture diagram implies exists.
Five Questions for Recovery
You don't need to redesign your entire architecture to start testing resilience. Begin by asking five questions:
- What fails first? Identify the control-plane dependencies your recovery path relies on: DNS, identity, routing, certificates, configuration, orchestration, or third-party services.
- What would the team actually observe? For each failure scenario, define the signal that would tell you something is wrong. A green application dashboard is not enough if the failure happens before requests reach the application.
- Can you execute the recovery path? Test the actual mechanism rather than reviewing it on a diagram. If failover depends on a DNS change, routing decision, or control-plane API, exercise that path under realistic conditions.
- Who owns recovery? The on-call engineer may be responsible for responding to an incident, but someone should also own keeping recovery mechanisms, runbooks, dependencies, and tests current.
- What changed after the last test? Record what failed, what was surprising, and what assumptions proved wrong. Resilience improves when these findings become engineering work rather than remaining incident notes.
The goal is not to test every possible failure. It is to regularly challenge the assumptions that make recovery possible and turn the results into changes in the system.
The Control-Pane Trap Nobody Draws on the Architecture Diagram
Most discussions about cloud failover focus on the data plane, specifically about where traffic goes, which region handles the request, and what happens when an instance dies. The control plane, the system that decides where traffic should flow, gets far less attention. The TLS 1.3 incident above is a direct result of that gap.
The data plane worked perfectly. The control plane broke.

Figure 1. A control-plane failure can redirect traffic away from a healthy region while the data plane continues operating normally. (Source: created by author).
You couldn't see any of this from inside the service. The service and the load balancer were healthy, but traffic simply wasn't arriving. Synthetic monitoring showed a diverging pattern that internal telemetry could not surface, because it only observed the data plane. The failure was visible only from outside, through monitoring infrastructure that was itself part of the control plane.
This is the structural problem with control-plane dependencies. They are invisible in normal operation. You don't think of CDN health checks as part of your service. You think of them as monitoring. However, monitoring that gates traffic is part of the data path, even though it lives in the control plane. When it breaks, your data plane is fine and your service is unreachable.
The pattern repeats across cloud services. AWS STS is a clear example. The global endpoint at sts.amazonaws.com is "hosted in a single AWS Region, US East (N. Virginia). Like other endpoints, it doesn't provide automatic failover to endpoints in other Regions", as the AWS documentation explicitly states. Many ostensibly global services carry a regional dependency that doesn't appear in standard architecture diagrams. If that region degrades, the global service can fail in ways that bypass regional redundancy entirely.
This points to something broader. Many multi-region architectures are only multi-region in the data plane. Their recovery assumptions still collapse onto a small number of shared control-plane dependencies. Moreover, those dependencies rarely appear on the architecture diagram until something breaks them.
There is a reason for this layering. The data plane is built to keep working when the control plane is impaired. An Amazon EC2 instance keeps running even if the EC2 control plane degrades. An in-flight AWS Lambda invocation will complete even if the Lambda control plane stalls. That is good. Although when you build failover on top of the control plane, as DNS-based failover via Route 53 health checks does, you inherit those failure modes whether you intended to or not. Your failover only works if the system that decides to fail over is itself working.
The question worth asking when designing failover is therefore not just "What happens if a region fails?" but "What happens if the thing that detects regional failure is itself degraded?" That surfaces dependencies that would otherwise stay invisible until a TLS version upgrade or a configuration drift makes them visible the hard way.
Afterwards, two alerting gaps were identified and filled. The CDN had moved traffic away from an entire region without triggering an alert and load balancers had dropped to zero inbound traffic just as silently. Both had been silent throughout. The failure was caught by an engineer manually watching metrics during the rollout. It also emerged that the CDN had been silently failing health checks for a week on the first batch of upgraded load balancers, while traffic continued to flow normally on the remaining TLS 1.2 targets. Nobody had known.
ARC Versus DNS-Based Failover: An Honest Comparison
If control-plane dependencies are an accepted risk, the question is what to do about them. The two mechanisms most teams evaluate are AWS Application Recovery Controller (ARC) and traditional DNS-based failover via Route 53 health checks. They sound similar, but they aren't.
DNS-based failover is the well-worn default. Route 53 probes each endpoint, marks it healthy or unhealthy and updates DNS records accordingly. They are simple, widely understood, and cheap. But if your organization genuinely cares about sub-minute recovery objectives, DNS failover quickly becomes uncomfortable. DNS propagation delays (i.e., TTLs, recursive resolvers, and clients that cache aggressively) all become part of your recovery window, whether you planned for them or not. Add the control-plane dependency described above, and the picture gets worse. If Route 53's health check infrastructure degrades, your failover will either not trigger or will trigger incorrectly.
ARC takes a different approach. Instead of DNS, it provides routing controls you flip explicitly when you decide to fail over, with continuous readiness checks so you know whether the failover target can actually accept load before you flip. The routing control plane is a cluster of five regional endpoints designed to remain operable during regional impairment. ARC doesn't remove complexity. Instead, it relocates complexity into a more explicit and dedicated operational model.
- Speed
DNS is bounded by health check interval, propagation delay, and client TTL for minutes in practice, but sometimes for longer. ARC operates in seconds. If your RTO is in seconds, DNS is not an acceptable choice. - Failure modes
DNS failover depends on Route 53's health check infrastructure’s accuracy; the TLS 1.3 incident is exactly that dependency misfiring. ARC's cluster of five regional endpoints indicates a routing control flip doesn't require any single control plane to be healthy. - Operational complexity
DNS is something every engineer already understands. ARC is operationally heavier. The routing controls and readiness checks require dedicated ownership. Runbooks need to stay current. In addition, on-call need to know how to execute a flip under pressure, not discover the runbook is wrong mid-incident. These incidents are examples of real overhead. - Cost
DNS is effectively free. ARC charges per-cluster and per-control; such charges are not expensive for a single critical service, but they add up. - Pre-failover confidence
ARC's readiness checks continuously validate that the target can accept traffic. With DNS, you find out at failover time.
Given the benefits of each tool, a tiered approach works well, with ARC for the small set of customer-facing critical paths, where control-plane dependency is unacceptable and DNS for everything else. Payment flows, authentication, and core booking paths, each warrants ARC. Internal tools and low-traffic APIs, however, do not.
Recovery Is a Capability, Not a Property
Recovery capability is not something organizations discover during an incident. By that point, whatever capability exists has already been built or neglected months earlier.
The difficult part of building recoverable systems is not just organizational: Resilience cannot be fully proved. High availability has benchmarks, including uptime, redundancy, failover timing, and replication lag. Resilience is also adversarial. Correlated failures, stale runbooks, control-plane coupling, and human coordination under stress, many of these issues only surface under conditions too expensive or too risky to reproduce reliably. Full regional failover under production load and recovery from a corrupted primary with real users waiting are not scenarios most teams can routinely exercise without accepting significant operational risk. This is the reason that performative resilience is so common, not because teams are careless, but because the bar for genuine proof is genuinely high.
The honest goal is not to guarantee recovery. In sufficiently complex systems, resilience is probabilistic, not provable. The realistic target is to improve confidence by reducing unknowns, rehearsing coordination, narrowing the blast radius, and shortening the gap between failure and detection. Architectural diagrams look resilient because the components are redundant. However, resilience is not determined by how the diagram looks under normal conditions. It is determined by whether people can actually recover a system under pressure, with degraded visibility, incomplete context, and dependencies that aren't behaving as expected.
Most organizations eventually discover that the recovery procedures they never exercised are the ones that fail first. That gap is not visible in the architecture. It becomes visible the first time the recovery path has to actually be used.