Linux Processes & Signals
A process is a running program: code, memory, open files, and at least one thread of execution, identified by a process ID (PID). Everything on a Linux system, from your shell to a database server to a container's entrypoint, is a process in a tree rooted at PID 1 (usually systemd). Signals are the kernel's asynchronous notifications to processes: "please terminate", "reload your config", "you've been interrupted".
Understanding processes and signals is essential for operating services. It's how you find what's consuming a machine's CPU, why a deploy hangs on shutdown, why a container ignores docker stop, and why there are <defunct> entries in ps.
TL;DR
- Each process has a PID, a parent (PPID), a user, a state, and resources. Threads share the process's memory.
- New processes are created by
fork()(copy) andexec()(replace with a new program). - Inspect with
ps,top/htop,pgrep,pstree, and/proc/<pid>/. - SIGTERM (15) asks politely and can be handled; SIGKILL (9) can't be caught. The kernel just ends the process.
- Graceful shutdown: handle SIGTERM, stop accepting work, finish in-flight requests, then exit.
- A zombie has exited but hasn't been reaped by its parent; PID 1 in containers must reap children.
Quick Example
A shell script that shuts down gracefully:
Core Concepts
Process Creation: fork and exec
A process creates another by fork(), which makes a near-identical copy (cheap thanks to copy-on-write memory). The child then usually calls exec() to replace itself with a different program. Your shell does exactly this for every command. The child inherits environment variables, open file descriptors, working directory, and user and group IDs, which is why leaked file descriptors and environment variables propagate to subprocesses.
Process States
Many processes stuck in D state point to storage problems, and they inflate the load average even when CPU is idle. See Linux performance troubleshooting.
Zombies and Orphans
When a process exits, it remains as a zombie until its parent calls wait() to collect the exit status. Zombies use no memory beyond a process table entry, but many of them signal a buggy parent. If a parent dies first, its children become orphans and are re-parented to PID 1 (or a subreaper), which reaps them.
In containers, your application often is PID 1. If it spawns children and never reaps them, zombies accumulate. Use a minimal init like tini (docker run --init) or make sure the app reaps children.
Priorities and Scheduling
The nice value (−20 highest priority to 19 lowest) influences CPU scheduling: nice -n 10 ./batch-job, renice +5 -p 1234. I/O priority uses ionice. For hard limits, use cgroups via systemd (CPUQuota=, MemoryMax=) or container resource limits.
Signals
Processes can handle (catch), ignore, or accept the default action for most signals. SIGKILL and SIGSTOP are always enforced by the kernel.
Graceful Shutdown
Service managers stop processes the same way. systemd, docker stop, and Kubernetes all send SIGTERM, wait a grace period (10–30 s by default), then send SIGKILL. A well-behaved service, on SIGTERM:
- Stops accepting new connections or jobs (marks itself unready).
- Finishes in-flight requests and flushes buffers.
- Closes connections and releases locks.
- Exits with code 0 before the grace period ends.
Services that ignore SIGTERM get killed mid-request, causing dropped connections and possibly corrupted work.
Job Control
In an interactive shell: command & runs in the background, Ctrl+Z stops the foreground job (SIGTSTP), bg and fg resume it, and jobs lists them. nohup or disown keep a job running after logout. For anything long-lived, use a service manager, or at least tmux/screen.
Best Practices
Handle SIGTERM in Every Service
Register a signal handler that triggers a graceful drain. Most frameworks have hooks (Node's process.on('SIGTERM'), Go's signal.NotifyContext, Python's signal module, Java's shutdown hooks). Test it: kill -TERM your process and watch in-flight requests complete.
Use exec in Wrapper Scripts
Entrypoint scripts should end with exec ./server, so the server replaces the shell and receives signals directly. Otherwise the shell is PID 1, the signal goes to the shell, and the app never hears it.
Reach for SIGKILL Last
kill -9 skips cleanup: temp files, locks, partial writes, and unflushed data remain. Try SIGTERM (and SIGINT) first, and wait before escalating.
Inspect /proc When Tools Fall Short
/proc/<pid>/ exposes command line, environment, limits, open files, memory maps, cgroup, and namespaces, which is invaluable for debugging what a process is actually doing and configured with.
Common Mistakes
Shell Form Entrypoints in Containers
See Dockerfile best practices.
Killing Parents to Clean Up Zombies Without Understanding
Zombies disappear when their parent reaps them or exits. Killing random processes to "clear zombies" can take down the service that owns them. Fix the parent's child-handling instead.
Assuming Children Die With the Parent
Killing a parent process doesn't automatically kill its children; they become orphans and keep running. Signal the whole process group (kill -TERM -<pgid>), or rely on systemd or container cgroups, which clean up everything in the unit.
FAQ
What's the difference between SIGTERM and SIGKILL?
SIGTERM is a request: the process can catch it, clean up, and exit gracefully (or even ignore it). SIGKILL is enforced by the kernel immediately; the process gets no chance to run any code. Always try SIGTERM first.
What is a zombie process and is it harmful?
A zombie is a process that has exited but whose parent hasn't yet collected its exit status. It uses almost no resources, but it occupies a PID. A few are harmless; thousands indicate a parent that never calls wait(), and can exhaust PIDs. Fix the parent, or run a proper init as PID 1.
Why won't a process die even with kill -9?
It's probably in uninterruptible sleep (D state), blocked in the kernel waiting for I/O, often from a hung NFS mount or a failing disk. The signal is delivered when the I/O returns. Investigate the storage issue. It may also be a zombie, which is already dead and waiting to be reaped.
What's the difference between a process and a thread?
A process has its own address space and resources. Threads are execution units within a process that share its memory and file descriptors. On Linux both are "tasks" scheduled by the kernel, and threads show up under /proc/<pid>/task/, with ps -L or top -H.
Related Topics
- Linux — The operating system overview
- systemd — Supervising service processes
- Linux Performance Troubleshooting — Diagnosing busy and stuck processes
- Kubernetes Workloads — SIGTERM and termination grace periods
- Dockerfile Best Practices — Exec-form entrypoints
- Bash — Traps and job control in scripts