AWS RDS
Amazon Relational Database Service (RDS) runs relational databases for you: PostgreSQL, MySQL, MariaDB, SQL Server, Oracle, and IBM Db2, plus Amazon Aurora, AWS's cloud-native MySQL- and PostgreSQL-compatible engine. AWS handles provisioning, OS and engine patching, automated backups, failover, and monitoring; you handle schema, queries, sizing, and configuration.
RDS is the default choice for relational data on AWS. It removes most operational toil while keeping a familiar database engine, so your application talks to it exactly as it would to a self-hosted PostgreSQL or MySQL. The trade-offs are less control over the host, some restricted superuser capabilities, and a price premium over self-managed EC2.
TL;DR
- RDS manages patching, backups, failover, and monitoring; you still own schema, indexes, and queries.
- Multi-AZ = high availability (synchronous standby, automatic failover). Read replicas = read scaling (asynchronous). They solve different problems.
- Automated backups enable point-in-time recovery within the retention window (up to 35 days).
- Aurora separates compute from a distributed storage layer — faster failover, up to 15 low-lag replicas, serverless option.
- Use parameter groups for engine config, RDS Proxy for connection pooling, IAM auth or Secrets Manager for credentials.
- Keep instances in private subnets, encrypt at rest, and enable deletion protection in production.
Quick Example
A production PostgreSQL instance in Terraform — private, encrypted, Multi-AZ, with backups and Performance Insights:
Add a read replica for reporting queries with a second resource that sets replicate_source_db = aws_db_instance.app.identifier.
Core Concepts
Engines: RDS vs Aurora
Multi-AZ vs Read Replicas
- Multi-AZ deployments keep a synchronous standby in another availability zone. On failure or maintenance, RDS flips the DNS endpoint to the standby. The standby doesn't serve reads (except in Multi-AZ cluster deployments with two readable standbys).
- Read replicas use asynchronous replication to serve read traffic and can live in other regions for disaster recovery. They can lag, so don't send read-after-write traffic to them without care. See Database Replication.
Backups and Recovery
Automated backups take daily snapshots plus transaction logs, allowing restore to any second within the retention period (1–35 days). Manual snapshots persist until deleted and can be copied across regions and accounts. Restores always create a new instance — plan how the application switches endpoints. See Database Backups.
Configuration: Parameter and Option Groups
You can't edit postgresql.conf directly. Parameter groups hold engine settings (work_mem, max_connections, log_min_duration_statement); some apply immediately, others need a reboot. Option groups enable engine features for Oracle, SQL Server, and MySQL.
Connections and RDS Proxy
Each database connection consumes memory, and serverless or heavily scaled apps can exhaust max_connections. RDS Proxy pools and shares connections, speeds up failover handling, and integrates with IAM and Secrets Manager. For PostgreSQL, see also PostgreSQL Connection Pooling.
Storage
gp3 offers baseline IOPS and throughput you can raise independently of size; io2 provides provisioned high IOPS with low latency for demanding workloads. Storage autoscaling grows volumes automatically up to a ceiling.
Monitoring
CloudWatch metrics (CPU, freeable memory, IOPS, replica lag), Enhanced Monitoring (OS-level metrics), and Performance Insights (database load by wait event and SQL statement) show where time goes.
Best Practices
Run Production Multi-AZ
A single-AZ instance means an AZ issue or even routine maintenance becomes downtime. Multi-AZ roughly doubles instance cost but is standard for production.
Keep It Private and Encrypted
Place instances in private subnets with no public access, allow connections only from application security groups, enable encryption at rest with KMS, and require TLS in transit.
Manage Credentials Automatically
Use manage_master_user_password (Secrets Manager with rotation) or IAM database authentication. Never hard-code database passwords. See Secrets Management.
Tune With Evidence
Turn on Performance Insights and slow query logging. Fix the top queries with indexes and rewrites before resizing the instance. See Query Optimization.
Test Restores and Failovers
Restore a snapshot to a new instance on a schedule and run a failover (reboot with failover) in staging to measure real recovery times.
Right-Size and Reserve
Use Graviton (r7g, m7g) instances for better price-performance, downsize idle instances, stop non-production instances out of hours, and buy Reserved Instances for steady workloads.
Common Mistakes
Treating Read Replicas as High Availability
Replicas are asynchronous and don't fail over automatically (outside Aurora). Use Multi-AZ for HA.
Public Accessibility "Just for Testing"
publicly_accessible = true with a broad security group exposes the database to internet scanning. Use a bastion, SSM Session Manager, or a VPN.
Running Out of Connections
Lambda functions or autoscaled containers opening their own connections can hit max_connections and cause errors. Use RDS Proxy or an application-side pool.
Ignoring Storage and Burst Credits
Older gp2 volumes and burstable t instances can run out of credits and slow down suddenly under sustained load.
Deleting Without a Final Snapshot
Disable skip_final_snapshot in production and keep deletion_protection on.
FAQ
What is AWS RDS?
Amazon RDS is a managed service for running relational databases on AWS. It automates provisioning, patching, backups, and failover for engines like PostgreSQL, MySQL, SQL Server, and Oracle, plus AWS's Aurora.
What's the difference between Multi-AZ and a read replica?
Multi-AZ keeps a synchronous standby for automatic failover and high availability. A read replica is an asynchronous copy used to scale reads or for cross-region disaster recovery.
Should I use Aurora or standard RDS?
Aurora gives faster failover, low-lag replicas, and serverless scaling, at a higher baseline cost. Standard RDS is simpler and often cheaper for small or steady workloads, and supports more engines.
How do backups work in RDS?
Automated backups combine daily snapshots with transaction logs, allowing point-in-time restore within your retention window. Manual snapshots are kept until you delete them. Restores create a new database instance.
RDS or a database on EC2?
Choose RDS unless you need OS-level access, unsupported extensions or versions, or have the team and scale to justify self-management. The operational savings usually outweigh the price premium.
Related Topics
- PostgreSQL — The most popular engine on RDS
- MySQL — The other common RDS engine
- Database Replication — How replicas and standbys work
- Google Cloud SQL — The equivalent service on GCP
- Database Backups — Backup strategies and PITR