Design Resilient Architectures
Resilient architectures stay available despite component, Availability Zone, or Region failures. This chapter covers high availability, decoupling, fault tolerance, and the AWS services that let workloads self-heal and recover.
High Availability Across Availability Zones
Spreading resources across multiple Availability Zones removes single points of failure from AZ-level outages. RDS Multi-AZ maintains a synchronous standby and fails over automatically, while Auto Scaling groups spanning several AZs replace unhealthy instances and keep capacity balanced. Elastic Load Balancers distribute traffic only to healthy targets across those zones. Designing for at least two AZs is a baseline for production availability.
Decoupling Components
Loose coupling isolates failures and smooths demand spikes. Amazon SQS queues buffer work between producers and consumers so a slow or failed downstream tier does not lose messages, and SNS fans out notifications to multiple subscribers. EventBridge routes events between services with rules. Decoupled tiers can scale and fail independently, which is central to resilient design.
Data Durability and Backup
Choose storage that matches durability needs and plan for recovery. S3 stores objects redundantly across multiple AZs and offers versioning to protect against accidental deletion or overwrite. EBS snapshots and RDS automated backups enable point-in-time recovery, and cross-Region copies protect against Regional loss. Define recovery point and recovery time objectives to guide the backup strategy.
Resilient Databases
Managed databases reduce the operational burden of achieving resilience. Aurora replicates six copies of data across three Availability Zones and fails over quickly to a replica, while DynamoDB is inherently multi-AZ and offers global tables for multi-Region replication. Read replicas scale reads but are asynchronous and not a substitute for Multi-AZ failover. Match the engine to the workload's consistency and availability requirements.
Multi-Region and DNS Failover
For the highest availability, extend architectures across Regions. Route 53 health checks with failover routing detect an unhealthy primary endpoint and direct users to a secondary Region automatically. Global Accelerator can also route around unhealthy Regional endpoints over the AWS backbone. Regularly test failover so recovery works when a real outage occurs.