Infrastructure Architecture
Infrastructure architecture is the set of decisions that determine how software actually runs: how it's networked, where it computes, where data lives, who can change what, how releases ship, and how you know it's healthy. These decisions shape reliability, security, and cost long before any feature is written — and they're expensive to change once production traffic depends on them.
Good infrastructure is boring on purpose: repeatable, owned, observable, and built from managed services where possible. This page walks through the building blocks, the choices within each, and the organizational practices — environments, tagging, cost, and recovery — that keep infrastructure healthy as it grows.
TL;DR
- Decide each layer deliberately: network, compute, data, identity, delivery, observability.
- Prefer managed services unless running it yourself is a differentiator.
- Define everything in infrastructure as code and change it through review.
- Separate environments with account/project boundaries, not just naming.
- Tag everything with owner, environment, and cost center from day one.
- Design recovery objectives (RTO/RPO) before you need them, then test them.
Quick Example
A typical small-to-medium production layout on a public cloud, expressed as the decisions that matter:
Each line is a decision someone made on purpose, captured in code, and reviewable.
Core Concepts
The Building Blocks
Choosing Compute
Choose per workload, not per company. A system can reasonably run APIs on containers, background processing on functions, and a legacy service on a VM.
Network Design
- Put only load balancers and gateways in public subnets; applications and databases stay private.
- Spread across at least two, preferably three, availability zones for high availability.
- Plan IP ranges so VPCs, on-premises networks, and future peers don't overlap.
- Use private endpoints for cloud services so traffic doesn't cross the public internet.
- Control egress with NAT gateways and, where required, egress firewalls.
See Cloud Networking for details.
Environments and Landing Zones
A landing zone is the baseline structure every workload lands in: an account or project hierarchy, identity integration, network topology, logging, and guardrail policies. Separate accounts (AWS), subscriptions (Azure), or projects (GCP) per environment give hard boundaries for permissions, quotas, and blast radius — far stronger than naming resources prod-* in a shared account.
A common pattern is dev → staging → prod, with staging built from the same code as production so it actually predicts production behavior.
Tagging and Ownership
Tags (labels) connect resources to owners, environments, cost centers, and data classifications. Without them you can't answer "who owns this?" during an incident or "what does this team cost?" at month end. Enforce required tags with policy (AWS SCPs/Config, Azure Policy, GCP Organization Policy) and in IaC modules.
Reliability and Recovery
Define, per service, a recovery time objective (how long it can be down) and a recovery point objective (how much data you can lose). Those numbers decide between single-zone, multi-zone, and multi-region designs, and between nightly backups and continuous replication. See Disaster Recovery Planning.
Best Practices
Managed Services First
Every self-hosted database, queue, or cluster is a system you must patch, back up, scale, and page someone for. Use managed services unless you have a concrete reason — cost at scale, a missing feature, or regulatory constraints.
Everything in Code, Everything Reviewed
Provision with Terraform, OpenTofu, or Pulumi. Changes go through pull requests with a plan output attached. Console changes are for emergencies and are reconciled back into code immediately.
Least Privilege by Default
Humans get SSO and short-lived elevated roles; workloads get their own identities with narrowly scoped permissions. No shared admin accounts, no long-lived access keys on laptops.
Build Guardrails, Not Gates
Policies that prevent public buckets, unencrypted volumes, or untagged resources let teams self-serve safely. Platform engineering turns these into golden-path templates.
Make Cost Visible Early
Set budgets and anomaly alerts per account, review cost by tag monthly, and delete idle resources automatically in non-production. See Cloud Costs.
Observe From Day One
Ship logs, metrics, and traces to a central place before the first production launch, with alerts tied to user-facing symptoms.
Common Mistakes
"Shared Everything" Environments
Dev, staging, and prod in one account with one set of credentials means a script run against the wrong environment can delete production. Use separate accounts or projects.
Databases With Public IPs
Exposing a database to the internet "temporarily" is a leading cause of breaches. Keep data stores private and reach them through bastions, VPNs, or private endpoints.
Single Availability Zone Production
One zone outage takes everything down. Multi-AZ is usually a small cost increase for a large reliability gain.
No Cost Governance Until the Bill Arrives
Forgotten GPU instances, unbounded log retention, and cross-region data transfer surprise teams every month. Tag, budget, and alert from the start.
Snowflake Servers
Hand-configured servers can't be rebuilt reliably. If you can't recreate it from code, you don't really control it.
FAQ
What is infrastructure architecture?
It's the design of the underlying systems software runs on — networks, compute, storage, identity, delivery pipelines, and observability — and the decisions about how they're organized, secured, and operated.
How do I choose between VMs, containers, and serverless?
Match the model to the workload. Containers suit most long-running web services; serverless suits spiky or event-driven work with short execution times; VMs suit legacy software or special OS requirements. Many systems combine all three.
What does "managed services first" mean?
Prefer cloud-provided services — managed databases, queues, load balancers — over running the same software yourself. The provider handles patching, backups, and scaling, which frees your team to work on the product.
How many environments do I need?
Most teams need at least development, staging, and production, isolated at the account or project level. Add ephemeral preview environments per pull request when the platform makes them cheap.
What is a landing zone?
A landing zone is a pre-configured cloud foundation — account structure, identity, networking, logging, and guardrail policies — that every new workload is deployed into so it starts secure and consistent.
Related Topics
- Infrastructure as Code — Defining all of this in reviewed code
- Cloud Computing — The platforms and managed services
- High Availability — Designing for failure across zones and regions
- System Design — Application architecture on top of the infrastructure
- Cloud Costs — Keeping the bill understandable