Design High-Performing Architectures
High performance, about 24 percent of SAA-C03, is the ability to use computing resources efficiently to meet demand as it changes, aligning with the Performance Efficiency pillar. This domain asks you to select purpose-built compute, storage, database, and networking services, to scale elastically, and to add caching and content delivery so users experience low latency at any load. The theme is picking the right tool for the access pattern rather than forcing one service to do everything. This chapter covers elastic and right-sized compute, storage selection, choosing databases, caching layers, and networking and content-delivery performance, always framed around when each option is the best fit.
Selecting and Scaling Compute
Matching compute to the workload and letting it scale automatically is the heart of performance efficiency. Amazon EC2 offers instance families tuned to different needs: general purpose (M and T) balance resources, compute optimized (C) suit CPU-bound work like batch processing and gaming servers, memory optimized (R and X) suit in-memory databases and large caches, storage optimized (I and D) suit high local IOPS, and accelerated computing (P, G, Inf) suit machine learning and graphics. Choosing the correct family prevents both bottlenecks and overprovisioning. Layer EC2 Auto Scaling on top so the fleet grows and shrinks with demand: use target tracking to hold a metric such as average CPU at a set value, step scaling for graduated responses, and scheduled scaling for predictable patterns, and enable predictive scaling to provision ahead of forecasted load. For workloads where you would rather not manage servers at all, serverless options often perform better with less effort: AWS Lambda runs event-driven functions that scale automatically from zero to thousands of concurrent executions and bills per millisecond, ideal for spiky or unpredictable traffic and short tasks, while AWS Fargate runs containers on ECS or EKS without managing the underlying instances. Choose containers on ECS or EKS for long-running services and portability, and Lambda for event-driven, bursty, or glue logic. To smooth demand and let compute scale independently, decouple tiers with queues as covered in the resilience domain. On the exam, when a workload is unpredictable or event-driven, favor serverless; when it is steady and long-running, size an appropriate EC2 family with Auto Scaling; and always let elasticity, not a fixed fleet, absorb changes in load so performance stays steady and you pay only for what you use.
Choosing the Right Storage
AWS offers distinct storage services, and performance questions hinge on matching the storage type to the access pattern. Amazon S3 is object storage for virtually unlimited, durable data accessed over HTTPS, ideal for static assets, data lakes, backups, and media; it is not a file system or a boot volume. Amazon EBS provides block storage attached to a single EC2 instance (with Multi-Attach as an exception for certain volumes) and is the choice for boot volumes and databases; pick gp3 general-purpose SSD for most workloads because you can provision IOPS and throughput independently of size, io2 Block Express provisioned IOPS SSD for the highest-performance, latency-sensitive databases, and st1 or sc1 HDD volumes for throughput-oriented or cold, sequential workloads. Instance store provides ephemeral, physically attached disks with very high IOPS for temporary data such as caches and scratch space, but data is lost when the instance stops, so never use it for anything that must persist. Amazon EFS is a fully managed, elastic NFS file system shared by many Linux instances across AZs, the right pick when multiple instances need concurrent shared file access. Amazon FSx offers managed file systems for specific ecosystems: FSx for Windows File Server for SMB and Active Directory workloads, and FSx for Lustre for high-performance computing and machine-learning workloads that need extreme throughput and can link to S3. To speed uploads of large objects over long distances, use S3 Transfer Acceleration, and use S3 multipart upload for large files. On the exam, map the requirement to the service: single-instance block storage to EBS, shared Linux file storage to EFS, Windows shared storage to FSx for Windows, massive parallel throughput to FSx for Lustre, object storage to S3, and disposable high-speed scratch to instance store.
Choosing Purpose-Built Databases
Selecting a database that fits the access pattern is one of the most tested performance skills, because forcing one engine to serve every use case creates bottlenecks. For relational workloads needing joins, transactions, and strong consistency, use Amazon RDS or, for higher throughput and cloud-native scaling, Amazon Aurora; Aurora Serverless v2 scales capacity automatically for variable relational load. For key-value and document workloads that need consistent single-digit-millisecond latency at any scale with virtually unlimited throughput, choose Amazon DynamoDB, a serverless NoSQL database that removes capacity planning with on-demand mode and scales seamlessly, making it the default answer for high-performance, massively scalable simple-query workloads. For in-memory data structures and ultra-low-latency access, use Amazon ElastiCache. For analytics and complex queries over large historical datasets, use Amazon Redshift, a columnar data warehouse built for OLAP, rather than a transactional database. Other purpose-built options appear in scenarios: Amazon Aurora and RDS for OLTP, Amazon Neptune for graph relationships, Amazon Timestream for time-series data, Amazon DocumentDB for MongoDB-compatible document workloads, Amazon Keyspaces for Cassandra, and Amazon OpenSearch Service for full-text search and log analytics. Read replicas add read scaling to RDS and Aurora, and DynamoDB global tables add multi-Region read locality. The exam pattern is to read the described access pattern and pick the matching engine: high write-and-read scale with simple keys points to DynamoDB, relational integrity points to Aurora or RDS, analytics over big data points to Redshift, search points to OpenSearch, and graph traversal points to Neptune. Combining a transactional store with a separate analytics store, feeding data from DynamoDB or RDS into Redshift or a data lake, is a common high-performance design that keeps each workload on the engine built for it.
Caching and Content Delivery
Caching places data closer to the consumer and offloads backends, which is one of the most effective ways to lower latency and increase throughput. For database and application caching, Amazon ElastiCache runs managed Redis or Memcached in memory: choose Redis (OSS) or Valkey when you need persistence, replication, pub/sub, sorting, or complex data structures, and Memcached for a simple, horizontally scalable cache; a common pattern is a cache-aside layer in front of RDS to serve hot queries in microseconds and shield the database from read pressure. For DynamoDB specifically, DynamoDB Accelerator (DAX) is a purpose-built in-memory cache that turns single-digit-millisecond reads into microsecond reads without application changes. To store user session state and keep the application tier stateless, use ElastiCache or DynamoDB. For content delivery, Amazon CloudFront is a global content delivery network that caches web content, APIs, and media at hundreds of edge locations near users, cutting latency and origin load, and it also terminates TLS, integrates with WAF, and can run lightweight logic at the edge with CloudFront Functions or Lambda@Edge. Use CloudFront for cacheable HTTP content, static websites fronting an S3 origin, and to accelerate dynamic content. For non-HTTP or latency-sensitive TCP and UDP traffic that cannot be cached, AWS Global Accelerator improves performance by routing over the AWS backbone to the nearest healthy endpoint. API Gateway can cache responses to reduce backend calls. The exam rewards adding the right cache at the right layer: put CloudFront at the edge for global content, ElastiCache or DAX at the data layer for hot reads, and API Gateway caching for repeated API responses. When a question describes repeated expensive reads hammering a database, the fix is a cache; when it describes global users fetching static or cacheable content, the fix is CloudFront.
Keep going: the full AWS Solutions Architect Associate (SAA-C03) guide covers every section of the exam. AWS Solutions Architect Associate (SAA-C03) — Complete Study Guide (2026) — PDF + EPUB, $14.99 · 14-day refund →

Practice stays free. The full AWS Solutions Architect Associate (SAA-C03) study guide is the material itself, taught start to finish — a downloadable PDF + EPUB you keep.