Linux Performance Troubleshooting
When a Linux server is "slow", the cause is almost always one of four resources: CPU, memory, disk I/O, or network, or a software limit such as file descriptors, connection pools, or locks. The hard part isn't the tools. It's working through them systematically instead of guessing, so you find the actual bottleneck in minutes rather than restarting things and hoping.
Methods like Brendan Gregg's USE method (Utilization, Saturation, Errors per resource) and a quick "first 60 seconds" checklist of standard commands give you a repeatable path. When those point to a hot process, profilers like perf and eBPF tools show exactly where time goes.
TL;DR
- Use the USE method: for each resource, check utilization, saturation (queueing), and errors.
- Load average counts runnable and uninterruptible (I/O-waiting) tasks. Compare it to CPU count, and don't read it as CPU usage alone.
- CPU:
top/htop,mpstat -P ALL,pidstat, watching us/sy/wa/st. - Memory:
free -h("available" matters),vmstat(si/so swap activity), and OOM killer messages indmesg. - Disk:
iostat -xz 1(%util, await, queue size); network:ss -s,sar -n DEV, retransmits. - Go deeper with
perf, flame graphs, and eBPF/bcc tools (biolatency,execsnoop,tcpretrans).
Quick Example
The "first 60 seconds" checklist, adapted from Netflix's performance team:
Example reading: vmstat shows wa at 40% and iostat shows a disk at 100% %util with 80 ms await, while CPU us is low. It's a disk bottleneck, not a CPU one, even though load average is high.
Core Concepts
The USE Method
For each resource, ask three questions:
It prevents the common mistake of stopping at the first odd-looking metric.
Load Average
The 1-, 5-, and 15-minute load averages count tasks that are running, runnable, or in uninterruptible sleep (usually waiting on disk or NFS). A load of 16 on a 16-core machine may be fine; on a 4-core machine it's saturated. High load with low CPU usage points to I/O waits or D-state processes. See Linux processes & signals.
CPU
In top or vmstat:
- us: user code (your application).
- sy: kernel time; high values suggest syscall-heavy work, context switching, or networking.
- wa: idle while waiting for I/O; points to disk.
- st (steal): time the hypervisor gave to other VMs, which means a noisy neighbor or an undersized cloud instance.
- One core at 100% while others idle: a single-threaded bottleneck (
mpstat -P ALL).
Container CPU limits add throttling: check nr_throttled in the cgroup's cpu.stat.
Memory
Linux uses free memory for page cache, so "free" is usually low, and that's healthy. Look at available in free -h. Trouble signs:
- Swapping (
vmstat si/sonon-zero and sustained): the working set exceeds RAM, and performance collapses. - OOM killer:
dmesg -T | grep -i "killed process". The kernel killed a process to free memory, or a cgroup (container) hit itsmemory.max. - Growing RSS over days: likely a memory leak. Check with
pidstat -rorsmem.
Disk I/O
iostat -xz 1:
%utilnear 100% means the device is busy all the time. (For SSDs and NVMe, which handle parallel requests, also check latency.)r_await/w_await: average latency in ms. A few ms is typical for SSDs; tens or hundreds indicates saturation or a failing device.aqu-sz: average queue length, the saturation signal.
Find which process is doing I/O with iotop or pidstat -d. Full filesystems (df -h) and exhausted inodes (df -i) also cause failures that look like performance problems.
Network
ss -sandss -tanp: connection counts and states. Thousands ofTIME_WAITorCLOSE_WAITsockets indicate connection churn or an app not closing sockets.- Retransmits (
sar -n ETCP,nstat): packet loss or congestion. - Interface errors and drops (
ip -s link). - Latency to dependencies:
mtr,curl -wtiming,ping. See TCP/IP.
Going Deeper: Profiling
When a process is clearly burning CPU, find out where:
Flame graphs visualize sampled stacks: wide bars are where time goes. For managed runtimes, use language-aware profilers (async-profiler for the JVM, py-spy for Python, pprof for Go, 0x or --prof for Node).
eBPF tools (bcc, bpftrace) observe the kernel safely in production:
Best Practices
Establish Baselines
Know what normal looks like: typical load, CPU split, disk latency, and connection counts for each role. Continuous metrics (node_exporter plus Prometheus) make "is this unusual?" answerable and let you look at history from before the incident.
Check Errors and Logs Early
dmesg and system logs often explain everything at once: OOM kills, filesystem errors, NIC resets, and kernel warnings. Look before diving into metrics.
Change One Thing at a Time
When tuning (sysctl settings, limits, instance size), change one variable, measure, and keep notes. Stacking changes makes it impossible to know what helped.
Consider the Application Layer
Many "server is slow" problems are application issues: exhausted connection pools, lock contention, slow queries, garbage collection pauses, or a downstream dependency. System metrics tell you which resource is stressed; application metrics and traces tell you why. See performance engineering.
Common Mistakes
Panicking About Low "Free" Memory
Only 612 MiB is "free", but 20 GiB is available: page cache is reclaimed instantly when needed. This machine is fine.
Reading Load Average as CPU Usage
High load with idle CPUs usually means processes blocked on I/O. Check wa, iostat, and processes in D state before adding CPU.
Restarting Before Capturing Evidence
A restart may make the problem disappear temporarily, and destroy the evidence. Grab top, vmstat, iostat, ss, dmesg, and thread dumps first (a few seconds), then mitigate.
FAQ
What's a "good" load average?
Relative to CPU count: consistently below the number of cores means CPUs aren't saturated. Above it means tasks are queueing, whether for CPU or for I/O. Compare the 1-, 5-, and 15-minute values to see whether a problem is growing or subsiding.
How do I find what's using all the disk I/O?
iotop -oPa (accumulated I/O per process), pidstat -d 1, or eBPF's biosnoop for per-I/O detail. Then check whether it's expected, such as backups, compaction, or logs, or runaway behavior.
What does steal time mean on a cloud VM?
The hypervisor ran other tenants' work on the physical CPU your VM wanted. Sustained steal above a few percent degrades performance. Common causes are burstable instances that exhausted CPU credits and oversubscribed hosts. Move to a larger or non-burstable instance type.
When should I use perf or eBPF instead of top?
When top shows which process is busy but not why. perf and CPU profilers show which functions consume CPU, and eBPF tools reveal kernel-level behavior (I/O latency distributions, scheduler delays, TCP retransmits) that standard tools can't.
Related Topics
- Linux — The operating system overview
- Linux Processes & Signals — Process states and inspection
- Profiling — Finding hot code paths
- eBPF — Safe, deep kernel observability
- Troubleshooting — General troubleshooting methodology
- Performance Engineering — Performance as a discipline