Infrastructure Architecture

Infrastructure architecture is the set of decisions that determine how software actually runs: how it's networked, where it computes, where data lives, who can change what, how releases ship, and how you know it's healthy. These decisions shape reliability, security, and cost long before any feature is written — and they're expensive to change once production traffic depends on them.

Good infrastructure is boring on purpose: repeatable, owned, observable, and built from managed services where possible. This page walks through the building blocks, the choices within each, and the organizational practices — environments, tagging, cost, and recovery — that keep infrastructure healthy as it grows.

TL;DR

Quick Example

A typical small-to-medium production layout on a public cloud, expressed as the decisions that matter:

Each line is a decision someone made on purpose, captured in code, and reviewable.

Core Concepts

The Building Blocks

Choosing Compute

Choose per workload, not per company. A system can reasonably run APIs on containers, background processing on functions, and a legacy service on a VM.

Network Design

See Cloud Networking for details.

Environments and Landing Zones

A landing zone is the baseline structure every workload lands in: an account or project hierarchy, identity integration, network topology, logging, and guardrail policies. Separate accounts (AWS), subscriptions (Azure), or projects (GCP) per environment give hard boundaries for permissions, quotas, and blast radius — far stronger than naming resources prod-* in a shared account.

A common pattern is dev → staging → prod, with staging built from the same code as production so it actually predicts production behavior.

Tagging and Ownership

Tags (labels) connect resources to owners, environments, cost centers, and data classifications. Without them you can't answer "who owns this?" during an incident or "what does this team cost?" at month end. Enforce required tags with policy (AWS SCPs/Config, Azure Policy, GCP Organization Policy) and in IaC modules.

Reliability and Recovery

Define, per service, a recovery time objective (how long it can be down) and a recovery point objective (how much data you can lose). Those numbers decide between single-zone, multi-zone, and multi-region designs, and between nightly backups and continuous replication. See Disaster Recovery Planning.

Best Practices

Managed Services First

Every self-hosted database, queue, or cluster is a system you must patch, back up, scale, and page someone for. Use managed services unless you have a concrete reason — cost at scale, a missing feature, or regulatory constraints.

Everything in Code, Everything Reviewed

Provision with Terraform, OpenTofu, or Pulumi. Changes go through pull requests with a plan output attached. Console changes are for emergencies and are reconciled back into code immediately.

Least Privilege by Default

Humans get SSO and short-lived elevated roles; workloads get their own identities with narrowly scoped permissions. No shared admin accounts, no long-lived access keys on laptops.

Build Guardrails, Not Gates

Policies that prevent public buckets, unencrypted volumes, or untagged resources let teams self-serve safely. Platform engineering turns these into golden-path templates.

Make Cost Visible Early

Set budgets and anomaly alerts per account, review cost by tag monthly, and delete idle resources automatically in non-production. See Cloud Costs.

Observe From Day One

Ship logs, metrics, and traces to a central place before the first production launch, with alerts tied to user-facing symptoms.

Common Mistakes

"Shared Everything" Environments

Dev, staging, and prod in one account with one set of credentials means a script run against the wrong environment can delete production. Use separate accounts or projects.

Databases With Public IPs

Exposing a database to the internet "temporarily" is a leading cause of breaches. Keep data stores private and reach them through bastions, VPNs, or private endpoints.

Single Availability Zone Production

One zone outage takes everything down. Multi-AZ is usually a small cost increase for a large reliability gain.

No Cost Governance Until the Bill Arrives

Forgotten GPU instances, unbounded log retention, and cross-region data transfer surprise teams every month. Tag, budget, and alert from the start.

Snowflake Servers

Hand-configured servers can't be rebuilt reliably. If you can't recreate it from code, you don't really control it.

FAQ

What is infrastructure architecture?

It's the design of the underlying systems software runs on — networks, compute, storage, identity, delivery pipelines, and observability — and the decisions about how they're organized, secured, and operated.

How do I choose between VMs, containers, and serverless?

Match the model to the workload. Containers suit most long-running web services; serverless suits spiky or event-driven work with short execution times; VMs suit legacy software or special OS requirements. Many systems combine all three.

What does "managed services first" mean?

Prefer cloud-provided services — managed databases, queues, load balancers — over running the same software yourself. The provider handles patching, backups, and scaling, which frees your team to work on the product.

How many environments do I need?

Most teams need at least development, staging, and production, isolated at the account or project level. Add ephemeral preview environments per pull request when the platform makes them cheap.

What is a landing zone?

A landing zone is a pre-configured cloud foundation — account structure, identity, networking, logging, and guardrail policies — that every new workload is deployed into so it starts secure and consistent.

Related Topics

References