Incident management concept (SRE): adapted ICS for software incidents. Severity levels (SEV1/2/3), Incident Commander (IC) coordinator role separate from Operations who fixes, Communications Lead handling status page, Scribe capturing timeline, SMEs on-demand. Detection via Prometheus SLO burn -> Alertmanager -> PagerDuty -> on-call IC. War room via Slack channel + Zoom + Incident.io orchestration. Customer comms via Statuspage (Investigating -> Identified -> Monitoring -> Resolved cadence every 30min) + Twitter + support macros. Mitigation toolkit: rollback, feature flag off, regional failover, manual circuit break — mitigate first, investigate later. Four scenarios: SEV-1 textbook response with rollback in 28min, statuspage 30-min comms cadence, follow-the-sun handoff between US and EU regions, anti-pattern of investigation > mitigation causing 45-min downtime. Includes 2 ADRs: rotating IC vs dedicated IC team, blameless culture enforcement.
Incident management creates a temporary command structure for uncertain, time-sensitive work. Google SRE separates command, operations, communications, and planning. The incident commander coordinates objectives and priorities; Operations changes the system; Communications maintains stakeholder updates; Planning tracks future work and handoffs. Smaller incidents may combine roles deliberately, but responsibilities must remain explicit.
Mitigation restores acceptable service and can precede root-cause analysis, yet it is not “act without evidence.” Changes remain scoped, reviewed when feasible, logged, reversible, and verified against user impact. A live incident document records times, decisions, owners and evidence so that status and later learning do not depend on memory.
Detect, declare, and assign roles. At 12
monitoring detects impact; at 12 the commander declares the incident and assigns explicit responsibilities.Mitigate and verify. Operations chooses rollback at 12
, completes it at 12, and resolves only after user-level verification at 12.Evidence-based customer communication. Communications reports confirmed impact, mitigation and the next update without speculating about root cause.
Acknowledged handoff. The outgoing planning lead transfers current state and receives explicit acknowledgement from the next shift.
Введите числа или выберите пресет