Linux Processes & Signals

A process is a running program: code, memory, open files, and at least one thread of execution, identified by a process ID (PID). Everything on a Linux system, from your shell to a database server to a container's entrypoint, is a process in a tree rooted at PID 1 (usually systemd). Signals are the kernel's asynchronous notifications to processes: "please terminate", "reload your config", "you've been interrupted".

Understanding processes and signals is essential for operating services. It's how you find what's consuming a machine's CPU, why a deploy hangs on shutdown, why a container ignores docker stop, and why there are <defunct> entries in ps.

TL;DR

Quick Example

A shell script that shuts down gracefully:

Core Concepts

Process Creation: fork and exec

A process creates another by fork(), which makes a near-identical copy (cheap thanks to copy-on-write memory). The child then usually calls exec() to replace itself with a different program. Your shell does exactly this for every command. The child inherits environment variables, open file descriptors, working directory, and user and group IDs, which is why leaked file descriptors and environment variables propagate to subprocesses.

Process States

Many processes stuck in D state point to storage problems, and they inflate the load average even when CPU is idle. See Linux performance troubleshooting.

Zombies and Orphans

When a process exits, it remains as a zombie until its parent calls wait() to collect the exit status. Zombies use no memory beyond a process table entry, but many of them signal a buggy parent. If a parent dies first, its children become orphans and are re-parented to PID 1 (or a subreaper), which reaps them.

In containers, your application often is PID 1. If it spawns children and never reaps them, zombies accumulate. Use a minimal init like tini (docker run --init) or make sure the app reaps children.

Priorities and Scheduling

The nice value (−20 highest priority to 19 lowest) influences CPU scheduling: nice -n 10 ./batch-job, renice +5 -p 1234. I/O priority uses ionice. For hard limits, use cgroups via systemd (CPUQuota=, MemoryMax=) or container resource limits.

Signals

Processes can handle (catch), ignore, or accept the default action for most signals. SIGKILL and SIGSTOP are always enforced by the kernel.

Graceful Shutdown

Service managers stop processes the same way. systemd, docker stop, and Kubernetes all send SIGTERM, wait a grace period (10–30 s by default), then send SIGKILL. A well-behaved service, on SIGTERM:

  1. Stops accepting new connections or jobs (marks itself unready).
  2. Finishes in-flight requests and flushes buffers.
  3. Closes connections and releases locks.
  4. Exits with code 0 before the grace period ends.

Services that ignore SIGTERM get killed mid-request, causing dropped connections and possibly corrupted work.

Job Control

In an interactive shell: command & runs in the background, Ctrl+Z stops the foreground job (SIGTSTP), bg and fg resume it, and jobs lists them. nohup or disown keep a job running after logout. For anything long-lived, use a service manager, or at least tmux/screen.

Best Practices

Handle SIGTERM in Every Service

Register a signal handler that triggers a graceful drain. Most frameworks have hooks (Node's process.on('SIGTERM'), Go's signal.NotifyContext, Python's signal module, Java's shutdown hooks). Test it: kill -TERM your process and watch in-flight requests complete.

Use exec in Wrapper Scripts

Entrypoint scripts should end with exec ./server, so the server replaces the shell and receives signals directly. Otherwise the shell is PID 1, the signal goes to the shell, and the app never hears it.

Reach for SIGKILL Last

kill -9 skips cleanup: temp files, locks, partial writes, and unflushed data remain. Try SIGTERM (and SIGINT) first, and wait before escalating.

Inspect /proc When Tools Fall Short

/proc/<pid>/ exposes command line, environment, limits, open files, memory maps, cgroup, and namespaces, which is invaluable for debugging what a process is actually doing and configured with.

Common Mistakes

Shell Form Entrypoints in Containers

See Dockerfile best practices.

Killing Parents to Clean Up Zombies Without Understanding

Zombies disappear when their parent reaps them or exits. Killing random processes to "clear zombies" can take down the service that owns them. Fix the parent's child-handling instead.

Assuming Children Die With the Parent

Killing a parent process doesn't automatically kill its children; they become orphans and keep running. Signal the whole process group (kill -TERM -<pgid>), or rely on systemd or container cgroups, which clean up everything in the unit.

FAQ

What's the difference between SIGTERM and SIGKILL?

SIGTERM is a request: the process can catch it, clean up, and exit gracefully (or even ignore it). SIGKILL is enforced by the kernel immediately; the process gets no chance to run any code. Always try SIGTERM first.

What is a zombie process and is it harmful?

A zombie is a process that has exited but whose parent hasn't yet collected its exit status. It uses almost no resources, but it occupies a PID. A few are harmless; thousands indicate a parent that never calls wait(), and can exhaust PIDs. Fix the parent, or run a proper init as PID 1.

Why won't a process die even with kill -9?

It's probably in uninterruptible sleep (D state), blocked in the kernel waiting for I/O, often from a hung NFS mount or a failing disk. The signal is delivered when the I/O returns. Investigate the storage issue. It may also be a zombie, which is already dead and waiting to be reaped.

What's the difference between a process and a thread?

A process has its own address space and resources. Threads are execution units within a process that share its memory and file descriptors. On Linux both are "tasks" scheduled by the kernel, and threads show up under /proc/<pid>/task/, with ps -L or top -H.

Related Topics

References