Incident Management

An incident is an unplanned interruption or reduction in the quality of a service. Incident management is the practiced ability to recognize impact, assemble the right people, restore service safely, and keep affected users informed while facts are incomplete.

TL;DR

Quick Example

Illustrative incident record; times and severity definitions must follow your service policy.

Core Concepts

Incident, Problem, and Change

Incident management restores disrupted service. Problem management investigates recurring or underlying causes. Change management controls modifications to the service. One outage may create records in all three workflows; an incident can close after validated restoration while the problem investigation remains open.

Severity and Priority

Severity describes impact; priority determines response order using impact and urgency. A single blocked payroll operator near a deadline can deserve a fast response even when few users are affected.

The First 15 Minutes

  1. Confirm the symptom and affected service.
  2. Estimate scope: who is blocked, where, and since when?
  3. Assign severity from impact and urgency—not executive volume.
  4. Name an incident commander, technical lead, and communicator.
  5. Open one timeline and one coordination channel.
  6. Prefer reversible mitigation before deep diagnosis.

The incident commander manages priorities and coordination. The technical lead directs investigation. The communicator translates verified facts into predictable updates. Separating these roles keeps troubleshooters focused.

Restore Before You Explain

Rollback a suspect change, fail over, disable a broken feature, shed nonessential load, or provide a manual workaround when it reduces harm. Preserve logs and evidence, but do not delay a safe restoration just to find the perfect root cause. For suspected compromise, involve the security response lead: containment, evidence preservation, and a trusted recovery environment may take priority over reconnecting a service. Availability recovery must not reopen an attacker's access.

Communication Template

We are investigating [user-visible impact] affecting [scope] since [time]. The team is [current action]. The next update will be at [time], even if there is no material change.

Avoid unsupported causes and optimistic recovery times. State impact, action, and next update.

After Service Returns

Validate from the user's perspective, drain backlogs, close temporary access, and monitor for recurrence. A post-incident review should explain contributing conditions, why defenses did not catch them, what helped recovery, and which actions have owners and dates. Measure detection time, acknowledgment time, restoration time, update reliability, recurrence, and overdue follow-ups.

Comparison

Best Practices

Make Handoffs Explicit

When responders change shifts, record the current impact, unsuccessful hypotheses, active changes, risks, and next update time. Have the incoming commander acknowledge ownership.

Check the Whole Recovery Path

A responding health endpoint is insufficient if login, writes, or queued work still fail. Include user confirmation and backlog checks in closure criteria.

Common Mistakes

Restoring Without Validation

Bad: Close the incident when a server restarts.

Correct: Check the affected workflow, error rate, dependent services, and backlog before declaring recovery.

Making Concurrent Untracked Changes

Bad: Let every responder modify production independently.

Correct: Coordinate changes through the technical lead and timestamp each intervention.

FAQ

Does every incident need three people?

No. A small incident may combine roles, but name who owns coordination and updates. Split roles as scope or cognitive load grows.

Must the root cause be known before closure?

No. Restore and validate service, then track remaining investigation separately with an owner.

How often should updates be sent?

Choose a cadence based on impact and audience. Announce the next update time and keep it even when the investigation has not changed.

Related Topics

References