Incident Management definition
Incident management is the process teams use to detect, respond to, resolve and learn from unplanned disruptions to a service, such as outages, security events or severe performance problems. It defines who is alerted, who leads the response, how stakeholders are informed and how follow-up actions prevent the same incident from happening again.
The incident management lifecycle
A good process is decided before anything breaks. During an outage, nobody should be debating who is in charge or where to talk. The lifecycle below is common across software teams, from small startups to large SRE organizations. Practice through game days keeps the process familiar.
- Detect: alerts from monitoring, or reports from users and support.
- Triage and declare: confirm impact and assign a severity level.
- Mobilize: page the on-call engineer and open a dedicated channel.
- Mitigate: restore service first, through rollback, failover or disabling a feature.
- Communicate: update stakeholders and a status page at regular intervals.
- Resolve: confirm the service is healthy and close the incident.
- Learn: hold a postmortem and track follow-up actions to completion.
Severity levels and roles
Severity levels set the response. A SEV1 might be a full outage or data breach affecting all customers, needing everyone available now. A SEV2 might be a major feature broken for many users, and a SEV3 a minor issue handled in working hours. Clear definitions prevent both overreaction and underreaction. Severity can be raised or lowered as the picture becomes clearer.
Larger incidents assign roles, borrowed from emergency services' incident command systems. The incident commander coordinates and makes decisions but does not debug. A communications lead updates customers and executives. Subject matter experts investigate and fix, and a scribe records the timeline. Separating these jobs keeps engineers focused and stakeholders informed.
Blameless postmortems
After significant incidents, teams write a postmortem covering the timeline, customer impact, contributing factors, what went well and specific action items with owners and dates. Blameless means the analysis assumes people acted reasonably with the information they had, and asks how the system, tools and processes allowed the failure. That honesty surfaces real causes, which blame would hide, and shared postmortems spread lessons across teams.
Tools and regulatory duties
Common tools include PagerDuty, incident.io, Rootly and FireHydrant for alerting, on-call schedules and incident workflows, Slack or Microsoft Teams for coordination, and Atlassian Statuspage or similar for customer updates. Teams track mean time to acknowledge and mean time to resolve to spot trends.
Security incidents carry legal duties. Under GDPR, notifiable personal data breaches must be reported to the supervisory authority within 72 hours of becoming aware of them, and India's CERT-In directions require certain cyber incidents to be reported within six hours. Incident plans should include legal and compliance contacts.
Example: a payment failure incident
At 19:05, alerts show checkout errors rising sharply. The on-call engineer declares a SEV1, an incident commander takes charge and the status page is updated within minutes. Traces point to a configuration change deployed at 18:50, which is rolled back, and payments recover by 19:30. The postmortem adds a validation check for that configuration and a canary step to deployments. Nexzem sets up on-call, runbooks and postmortem practices for the systems it builds and supports.