Kubernetes Storage

Containers have ephemeral filesystems: when a container restarts, anything it wrote is gone. Kubernetes storage gives Pods data that outlives a container, a Pod, or even a node. It does that through a layer of abstractions that separate what an app needs ("20 GiB, read-write, fast") from how the cluster provides it (an AWS EBS volume, a Ceph block device, an NFS share).

Getting storage right matters most for stateful workloads such as databases, queues, and search indexes. There, a wrong access mode or reclaim policy can mean a stuck rollout or lost data.

TL;DR

Quick Example

A StorageClass, and a StatefulSet that dynamically claims one volume per replica:

When postgres-0 is scheduled, the CSI driver creates a 50 GiB encrypted gp3 volume in the same zone and attaches it. If the Pod moves, the volume follows it.

Core Concepts

Volume Types

PV, PVC, and StorageClass

The three objects divide responsibility:

A PVC and PV bind one-to-one. The Pod references the PVC by name and never needs to know which cloud or array is behind it.

Access Modes

Reclaim Policy and Binding Mode

reclaimPolicy decides what happens to the PV when its PVC is deleted: Delete (the default for dynamic provisioning) destroys the underlying disk, while Retain keeps it for manual recovery. volumeBindingMode: WaitForFirstConsumer delays provisioning until a Pod is scheduled, so the volume is created in the same availability zone as the Pod. Immediate can create a volume in a zone the Pod can't reach.

CSI Drivers and Snapshots

The Container Storage Interface is the plugin standard for storage vendors. CSI drivers provision, attach, mount, resize, and snapshot volumes. With the snapshot CRDs installed, a VolumeSnapshot captures a point-in-time copy that can seed a new PVC. That's useful for clones and fast restores, but it doesn't replace off-cluster backups.

Stateful Workloads in Practice

Running databases on Kubernetes is viable today, but storage is where it gets hard:

Best Practices

Use WaitForFirstConsumer

For any zonal storage, delayed binding avoids the classic "volume in us-east-1a, Pod scheduled in us-east-1b" deadlock.

Retain Important Data

Create a separate StorageClass with reclaimPolicy: Retain for production databases. Deleting a namespace or a Helm release then no longer deletes the disk underneath.

Enable Volume Expansion

Set allowVolumeExpansion: true so you can grow a PVC by editing its size instead of migrating data. Volumes can grow but never shrink, so start modestly.

Keep Scratch Data in emptyDir

Caches, temp files, and build artifacts don't need persistent volumes. Set sizeLimit on emptyDir so a runaway process can't fill the node's disk and trigger evictions.

Common Mistakes

Scaling a Deployment That Uses an RWO Volume

Either use an RWX filesystem, give each replica its own volume with a StatefulSet, or better, move uploads to object storage such as S3.

Using hostPath for Application Data

hostPath ties data to one node, bypasses scheduling, and is a common container escape vector. Reserve it for node agents (DaemonSets) that genuinely need host files.

Assuming Snapshots Are Backups

A snapshot deleted along with the cluster, the account, or the region is not a backup. Follow the 3-2-1 rule from backup strategy.

FAQ

What's the difference between a PV and a PVC?

A PersistentVolume is the actual storage resource in the cluster. A PersistentVolumeClaim is a namespaced request for storage that gets bound to a matching PV. Apps reference claims; admins or provisioners supply volumes.

Can I share a volume between Pods on different nodes?

Only with an access mode of ReadWriteMany, which requires a shared filesystem such as NFS, Amazon EFS, Azure Files, or CephFS. Block storage (EBS, Persistent Disk) attaches to one node at a time. For shared app data, object storage is usually simpler and scales better.

Should I run my database on Kubernetes?

It's reasonable if you have a solid operator, fast storage, and people who understand both the database and Kubernetes. Many teams still choose a managed database (RDS, Cloud SQL, Neon) and keep Kubernetes stateless, trading some cost for much less operational risk.

How do I resize a volume?

If the StorageClass allows expansion, edit the PVC's spec.resources.requests.storage to a larger value. The CSI driver grows the disk and, for most filesystems, the filesystem online. For StatefulSets you edit each PVC; the volumeClaimTemplates field is immutable.

Related Topics

References