KnowBite KnowBite

Engineering Leadership · · 1 min read

Article: High Availability Is Not Resilience: Why Cloud Systems Fail When It Matters Most

A routine TLS 1.3 upgrade silently broke Route 53 health checks, causing a CDN to stop routing traffic to a healthy region while internal dashboards showed nothing wrong. This article examines why HA and resilience are different problems, how control-plane dependencies create invisible failure modes, and why recovery capability erodes without explicit ownership. By Alexey Golev

Read original on InfoQ