Reliability and Business Continuity
This domain covers scaling, high availability, backup, and recovery so workloads survive failures and demand changes. Focus on Auto Scaling, multi-AZ designs, and backup strategy.
Elasticity with Auto Scaling
EC2 Auto Scaling groups launch and terminate instances to match demand across multiple Availability Zones. Target tracking, step, and scheduled policies adjust capacity, while health checks replace unhealthy instances automatically. Combine with an Elastic Load Balancer so traffic only reaches healthy targets.
High availability across AZs
Spreading resources across at least two Availability Zones protects against a single-AZ failure. RDS Multi-AZ maintains a synchronous standby with automatic failover, and load balancers distribute traffic across zones. Design so the loss of one AZ does not take down the workload.
Backup and recovery
AWS Backup centrally schedules and enforces backups across EBS, RDS, DynamoDB, and more with retention and vault policies. EBS snapshots are incremental and stored in S3, and point-in-time recovery protects databases. Test restores regularly so recovery objectives are actually met.
Recovery objectives and DR
RPO defines acceptable data loss and RTO defines acceptable downtime. Disaster-recovery strategies range from backup and restore, through pilot light and warm standby, to multi-site active/active, trading cost against faster recovery. Choose the pattern that meets the business RTO and RPO.