Database selection is one of the most important architecture skills on the AWS SAA-C03 exam because AWS offers relational and purpose-built services with very different data models, scaling behavior, latency, operational effort, and resilience characteristics. The right answer starts with the workload’s access pattern, not with a favorite database engine.
A useful decision sequence is to identify the data model, transaction requirements, query shape, scale, latency target, consistency needs, availability target, and operational constraints. Only then should the architect choose among Amazon RDS, Aurora, DynamoDB, DocumentDB, Neptune, ElastiCache, MemoryDB, Timestream, and other services.
Many real applications use more than one database type. That is not automatically overengineering if each store serves a clear access pattern.
Choose relational databases when relationships and transactions dominate
Amazon RDS is a managed relational database service that supports familiar engines such as PostgreSQL, MySQL, MariaDB, Oracle, SQL Server, and Db2. It is a strong fit when applications rely on SQL, joins, relational constraints, and transactional behavior.
RDS manages tasks such as backups, software patching, monitoring integration, and deployment options for high availability. Multi-AZ deployments improve availability, while read replicas can scale read workloads depending on the engine.
The Amazon RDS architecture is especially relevant when a scenario asks how to move a conventional relational application to AWS without taking on database operating-system management.
Use Aurora when the relational model fits but cloud-native scale matters
Amazon Aurora is compatible with MySQL and PostgreSQL ecosystems while using a distributed storage architecture designed for high availability and performance. It can separate writer and reader roles, add replicas for read scaling, and use serverless capacity options for certain workload shapes.
Aurora Global Database can support multi-Region read and disaster-recovery requirements, while newer Aurora variants address additional distributed SQL and scaling needs. The choice should be based on specific requirements rather than assuming Aurora is always better than standard RDS.
If the application depends on an engine feature or commercial database capability that Aurora does not provide, regular RDS may be the more compatible and lower-risk option.
Choose DynamoDB for key-value access at massive scale
Amazon DynamoDB is a fully managed serverless NoSQL database optimized for key-value and document access with single-digit millisecond performance at scale. It fits sessions, carts, metadata, state, gaming, IoT, and other workloads whose access patterns can be designed around partition and sort keys.
DynamoDB should not be selected simply because an application needs high scale. The data model matters. Complex ad hoc joins and relational queries are not its strength. The application must know how it will read data so the table and indexes can be designed around those patterns.
On-demand capacity can simplify unpredictable workloads, while provisioned capacity with auto scaling can fit more predictable traffic. Global tables add multi-Region active-active capabilities for eligible use cases.
Use document databases when the document is the natural aggregate
Amazon DocumentDB stores JSON-like documents and provides MongoDB compatibility for supported workloads. It can fit catalogs, content systems, profiles, and other applications where related fields naturally live together as documents and schemas change more flexibly than traditional tables.
Compatibility should be validated when migrating a MongoDB application because “compatible” does not mean every feature behaves identically. Query patterns, indexes, drivers, and unsupported features need testing.
Document storage can reduce impedance between application objects and database representation, but it is not a substitute for relational integrity when the business problem is fundamentally relational.
Use graph databases for relationship traversal
Amazon Neptune is designed for graph use cases such as social relationships, fraud networks, knowledge graphs, identity relationships, and recommendation systems. The value comes from traversing and evaluating connections efficiently.
A relational database can represent graphs, but complex multi-hop relationship queries may become cumbersome and expensive. A graph model is useful when relationships are first-class data rather than occasional joins.
SAA-C03 scenarios often signal Neptune through language about highly connected datasets, paths, networks of relationships, or graph queries.
Use caching to avoid repeatedly paying database latency
Amazon ElastiCache provides managed in-memory caching using engines such as Valkey, Redis OSS, or Memcached depending on the deployment option. It is useful when applications repeatedly read the same hot data and the authoritative database does not need to process every request.
Caching can improve latency and reduce load, but introduces consistency and invalidation decisions. The application must define what happens when cache content is stale or unavailable.
Amazon MemoryDB can be used when the workload needs an in-memory data model with durability as a primary database rather than a disposable cache. The design intent separates it from ordinary caching.
Choose time-series and wide-column stores when their models fit
Amazon Timestream is designed for time-series data such as application metrics, IoT telemetry, and operational measurements where time is a central dimension. Amazon Keyspaces provides a managed Cassandra-compatible wide-column model for workloads that need that access pattern.
Purpose-built databases exist because one database architecture cannot optimize every combination of latency, query, scale, and relationship. The challenge is to avoid using a specialized database when the team does not need its strengths.
Every additional data technology adds operational knowledge, monitoring, backup, IAM, and development requirements. Use polyglot persistence when the benefit is clear.
Availability and recovery options differ by service
Database selection includes failure behavior. RDS Multi-AZ, Aurora distributed storage and replicas, DynamoDB regional architecture, global tables, backups, point-in-time recovery, and cross-Region features all provide different recovery characteristics.
Do not assume a read replica is the same as a standby, or that multi-Region replication eliminates the need for backups. Logical corruption and accidental deletion can propagate through replication.
The high availability and fault-tolerance distinction helps clarify whether the requirement is automatic local failover, disaster recovery, or uninterrupted service across a larger failure domain.
Performance should be evaluated from the query pattern
Relational performance can depend on indexes, join complexity, connection handling, storage, and read scaling. DynamoDB performance depends heavily on key design and request distribution. Graph databases depend on traversal shape. Caches depend on hit rate and object size.
Architects should use representative load tests rather than assume a service will meet latency targets because its marketing category sounds appropriate. The database that fits the data model still needs the correct capacity and schema design.
For analytics, consider whether the operational database should serve analytical queries at all. Moving analytical work to a purpose-built analytics platform can protect transactional performance and lower coupling.
Cost follows architecture, not just database price
Compare infrastructure price with operational effort, scaling behavior, licensing, availability replicas, backup storage, data transfer, and engineering complexity. A managed service that costs more per hour may cost less overall if it removes patching and manual failover.
Conversely, a specialized database can be unnecessarily expensive if the workload could be served by an existing relational platform with acceptable performance. The AWS architecture certifications emphasizes these whole-system tradeoffs.
For SAA-C03, translate scenario language into data characteristics. Transactions and joins point toward relational databases. Predictable key access at massive scale points toward DynamoDB. Connected relationships suggest Neptune. Repeated hot reads suggest caching. Time-ordered telemetry suggests Timestream. Choosing from the access pattern is more reliable than memorizing service slogans.
Migration compatibility can outweigh theoretical service fit
A greenfield application can choose the ideal data model freely; a migrated application may be constrained by drivers, SQL dialect, stored procedures, licensing, or vendor support. Moving an Oracle or SQL Server application to Amazon RDS with the same engine can reduce operational effort without forcing an immediate application rewrite.
Engine conversion can still be worthwhile, but it should be treated as modernization rather than a transparent infrastructure move. AWS Database Migration Service can move data, while schema-conversion tooling and application testing may be needed when the source and target engines differ.
The safest migration answer often preserves compatibility first, then modernizes after the workload is stable. That sequence is especially relevant when the stated requirement prioritizes migration speed or low application change.
Connection management can become the bottleneck in serverless systems
Relational databases can be overwhelmed by thousands of short-lived connections even when CPU and storage are healthy. Serverless and highly elastic compute can create connection bursts that traditional application-server pools did not.
Amazon RDS Proxy can pool and share database connections for supported RDS and Aurora engines, helping applications scale connection demand more gracefully and improving resilience during some database failovers. It does not change the underlying data model; it addresses connection management.
This is a good example of architecture refinement after database selection. First choose the database that fits the data, then solve the access pattern around it with caching, proxies, replicas, or indexing where needed.
Backup and restore behavior should be tested for the chosen engine. Automated backups, snapshots, point-in-time recovery, and cross-Region copies have different recovery times and operational steps. The database that fits normal traffic but cannot meet the recovery objective is not the right architecture.
Likewise, verify service quotas and scaling limits before relying on a theoretical growth path. Database architecture should have a credible route from current load to expected peak without an emergency engine migration.