Kubernetes Storage
Containers have ephemeral filesystems: when a container restarts, anything it wrote is gone. Kubernetes storage gives Pods data that outlives a container, a Pod, or even a node. It does that through a layer of abstractions that separate what an app needs ("20 GiB, read-write, fast") from how the cluster provides it (an AWS EBS volume, a Ceph block device, an NFS share).
Getting storage right matters most for stateful workloads such as databases, queues, and search indexes. There, a wrong access mode or reclaim policy can mean a stuck rollout or lost data.
TL;DR
- Ephemeral volumes (
emptyDir,configMap,secret) live and die with the Pod. - A PersistentVolume (PV) is a piece of real storage; a PersistentVolumeClaim (PVC) is a request for storage that binds to a PV.
- A StorageClass describes a kind of storage (SSD, replicated, encrypted) and dynamically provisions PVs when a PVC asks for one.
- CSI drivers are the plugins that talk to actual storage backends (EBS, Persistent Disk, Azure Disk, Ceph, Longhorn).
- Access modes:
ReadWriteOnce(one node),ReadWriteOncePod,ReadOnlyMany,ReadWriteMany(many nodes; needs a shared filesystem). - Set
reclaimPolicy: Retainfor data you can't afford to lose, and use VolumeSnapshots plus real backups.
Quick Example
A StorageClass, and a StatefulSet that dynamically claims one volume per replica:
When postgres-0 is scheduled, the CSI driver creates a 50 GiB encrypted gp3 volume in the same zone and attaches it. If the Pod moves, the volume follows it.
Core Concepts
Volume Types
PV, PVC, and StorageClass
The three objects divide responsibility:
- The PVC is namespaced and owned by the app: "I need 50 GiB of
fast-ssd, ReadWriteOnce." - The StorageClass is cluster-scoped and owned by the platform team: which provisioner, which parameters, which reclaim policy.
- The PV is the actual volume. With dynamic provisioning it's created automatically to satisfy a PVC; with static provisioning an admin creates PVs ahead of time.
A PVC and PV bind one-to-one. The Pod references the PVC by name and never needs to know which cloud or array is behind it.
Access Modes
- ReadWriteOnce (RWO) — mounted read-write by Pods on one node. Block volumes (EBS, PD, Azure Disk) are RWO.
- ReadWriteOncePod (RWOP) — only a single Pod may mount it; stricter than RWO.
- ReadOnlyMany (ROX) — many nodes, read-only.
- ReadWriteMany (RWX) — many nodes, read-write. Requires a shared filesystem: NFS, EFS, Azure Files, CephFS.
Reclaim Policy and Binding Mode
reclaimPolicy decides what happens to the PV when its PVC is deleted: Delete (the default for dynamic provisioning) destroys the underlying disk, while Retain keeps it for manual recovery. volumeBindingMode: WaitForFirstConsumer delays provisioning until a Pod is scheduled, so the volume is created in the same availability zone as the Pod. Immediate can create a volume in a zone the Pod can't reach.
CSI Drivers and Snapshots
The Container Storage Interface is the plugin standard for storage vendors. CSI drivers provision, attach, mount, resize, and snapshot volumes. With the snapshot CRDs installed, a VolumeSnapshot captures a point-in-time copy that can seed a new PVC. That's useful for clones and fast restores, but it doesn't replace off-cluster backups.
Stateful Workloads in Practice
Running databases on Kubernetes is viable today, but storage is where it gets hard:
- Zones: an RWO block volume lives in one zone. If that zone fails, the Pod can't move until the volume is available again. Real HA comes from database-level replication across zones, not from the volume.
- Operators: mature operators (CloudNativePG, Strimzi for Kafka, the Elastic operator) handle failover, backups, and upgrades. Prefer them to hand-written StatefulSets.
- Backups: snapshots live in the same cloud account and region. Ship logical or physical backups off-cluster too; see database backups.
Best Practices
Use WaitForFirstConsumer
For any zonal storage, delayed binding avoids the classic "volume in us-east-1a, Pod scheduled in us-east-1b" deadlock.
Retain Important Data
Create a separate StorageClass with reclaimPolicy: Retain for production databases. Deleting a namespace or a Helm release then no longer deletes the disk underneath.
Enable Volume Expansion
Set allowVolumeExpansion: true so you can grow a PVC by editing its size instead of migrating data. Volumes can grow but never shrink, so start modestly.
Keep Scratch Data in emptyDir
Caches, temp files, and build artifacts don't need persistent volumes. Set sizeLimit on emptyDir so a runaway process can't fill the node's disk and trigger evictions.
Common Mistakes
Scaling a Deployment That Uses an RWO Volume
Either use an RWX filesystem, give each replica its own volume with a StatefulSet, or better, move uploads to object storage such as S3.
Using hostPath for Application Data
hostPath ties data to one node, bypasses scheduling, and is a common container escape vector. Reserve it for node agents (DaemonSets) that genuinely need host files.
Assuming Snapshots Are Backups
A snapshot deleted along with the cluster, the account, or the region is not a backup. Follow the 3-2-1 rule from backup strategy.
FAQ
What's the difference between a PV and a PVC?
A PersistentVolume is the actual storage resource in the cluster. A PersistentVolumeClaim is a namespaced request for storage that gets bound to a matching PV. Apps reference claims; admins or provisioners supply volumes.
Can I share a volume between Pods on different nodes?
Only with an access mode of ReadWriteMany, which requires a shared filesystem such as NFS, Amazon EFS, Azure Files, or CephFS. Block storage (EBS, Persistent Disk) attaches to one node at a time. For shared app data, object storage is usually simpler and scales better.
Should I run my database on Kubernetes?
It's reasonable if you have a solid operator, fast storage, and people who understand both the database and Kubernetes. Many teams still choose a managed database (RDS, Cloud SQL, Neon) and keep Kubernetes stateless, trading some cost for much less operational risk.
How do I resize a volume?
If the StorageClass allows expansion, edit the PVC's spec.resources.requests.storage to a larger value. The CSI driver grows the disk and, for most filesystems, the filesystem online. For StatefulSets you edit each PVC; the volumeClaimTemplates field is immutable.
Related Topics
- Kubernetes — The platform overview
- Kubernetes Workloads — StatefulSets and volume claim templates
- Database Backups — Protecting data beyond snapshots
- Database Replication — Real high availability for stateful systems
- AWS S3 — Object storage instead of shared volumes