Linux Performance Troubleshooting

When a Linux server is "slow", the cause is almost always one of four resources: CPU, memory, disk I/O, or network, or a software limit such as file descriptors, connection pools, or locks. The hard part isn't the tools. It's working through them systematically instead of guessing, so you find the actual bottleneck in minutes rather than restarting things and hoping.

Methods like Brendan Gregg's USE method (Utilization, Saturation, Errors per resource) and a quick "first 60 seconds" checklist of standard commands give you a repeatable path. When those point to a hot process, profilers like perf and eBPF tools show exactly where time goes.

TL;DR

Quick Example

The "first 60 seconds" checklist, adapted from Netflix's performance team:

Example reading: vmstat shows wa at 40% and iostat shows a disk at 100% %util with 80 ms await, while CPU us is low. It's a disk bottleneck, not a CPU one, even though load average is high.

Core Concepts

The USE Method

For each resource, ask three questions:

It prevents the common mistake of stopping at the first odd-looking metric.

Load Average

The 1-, 5-, and 15-minute load averages count tasks that are running, runnable, or in uninterruptible sleep (usually waiting on disk or NFS). A load of 16 on a 16-core machine may be fine; on a 4-core machine it's saturated. High load with low CPU usage points to I/O waits or D-state processes. See Linux processes & signals.

CPU

In top or vmstat:

Container CPU limits add throttling: check nr_throttled in the cgroup's cpu.stat.

Memory

Linux uses free memory for page cache, so "free" is usually low, and that's healthy. Look at available in free -h. Trouble signs:

Disk I/O

iostat -xz 1:

Find which process is doing I/O with iotop or pidstat -d. Full filesystems (df -h) and exhausted inodes (df -i) also cause failures that look like performance problems.

Network

Going Deeper: Profiling

When a process is clearly burning CPU, find out where:

Flame graphs visualize sampled stacks: wide bars are where time goes. For managed runtimes, use language-aware profilers (async-profiler for the JVM, py-spy for Python, pprof for Go, 0x or --prof for Node).

eBPF tools (bcc, bpftrace) observe the kernel safely in production:

See profiling and eBPF.

Best Practices

Establish Baselines

Know what normal looks like: typical load, CPU split, disk latency, and connection counts for each role. Continuous metrics (node_exporter plus Prometheus) make "is this unusual?" answerable and let you look at history from before the incident.

Check Errors and Logs Early

dmesg and system logs often explain everything at once: OOM kills, filesystem errors, NIC resets, and kernel warnings. Look before diving into metrics.

Change One Thing at a Time

When tuning (sysctl settings, limits, instance size), change one variable, measure, and keep notes. Stacking changes makes it impossible to know what helped.

Consider the Application Layer

Many "server is slow" problems are application issues: exhausted connection pools, lock contention, slow queries, garbage collection pauses, or a downstream dependency. System metrics tell you which resource is stressed; application metrics and traces tell you why. See performance engineering.

Common Mistakes

Panicking About Low "Free" Memory

Only 612 MiB is "free", but 20 GiB is available: page cache is reclaimed instantly when needed. This machine is fine.

Reading Load Average as CPU Usage

High load with idle CPUs usually means processes blocked on I/O. Check wa, iostat, and processes in D state before adding CPU.

Restarting Before Capturing Evidence

A restart may make the problem disappear temporarily, and destroy the evidence. Grab top, vmstat, iostat, ss, dmesg, and thread dumps first (a few seconds), then mitigate.

FAQ

What's a "good" load average?

Relative to CPU count: consistently below the number of cores means CPUs aren't saturated. Above it means tasks are queueing, whether for CPU or for I/O. Compare the 1-, 5-, and 15-minute values to see whether a problem is growing or subsiding.

How do I find what's using all the disk I/O?

iotop -oPa (accumulated I/O per process), pidstat -d 1, or eBPF's biosnoop for per-I/O detail. Then check whether it's expected, such as backups, compaction, or logs, or runaway behavior.

What does steal time mean on a cloud VM?

The hypervisor ran other tenants' work on the physical CPU your VM wanted. Sustained steal above a few percent degrades performance. Common causes are burstable instances that exhausted CPU credits and oversubscribed hosts. Move to a larger or non-burstable instance type.

When should I use perf or eBPF instead of top?

When top shows which process is busy but not why. perf and CPU profilers show which functions consume CPU, and eBPF tools reveal kernel-level behavior (I/O latency distributions, scheduler delays, TCP retransmits) that standard tools can't.

Related Topics

References