During a serious outage, clarity buys time. Teams need a first message that goes out fast, follow‑ups on a predictable cadence, and a crisp resolution note that closes the loop. The templates below are field‑tested, short, and specific. Fill the variables and send.
Why incident comms fail (and what to fix first)
- Ambiguous scope. Say exactly what is degraded and who is affected.
- No owner. Name the incident commander and the comms owner every time.
- Slipping cadence. Promise the next update time and meet it.
- Speculation. Report facts and current mitigation only. Keep guesses out.
Templates by severity
Choose SEV1 for critical customer‑impacting outages; SEV2 for serious partial impact. Edit bracketed variables and remove any lines that do not apply.
SEV1 initial (email)
Subject: [SEV1] [Service/component] outage — next update [HH:MM TZ]
Impact: [Who is affected and how].
Start: [HH:MM TZ, YYYY‑MM‑DD].
Scope: [Systems/regions].
Mitigation: [What is being done now].
Owner: [Incident commander name].
Next update: [HH:MM TZ] or sooner if material change.
Status page: [link]
SEV1 initial (Slack)
SEV1 — [Service/component]
Impact: [who/how]
Start: [HH:MM TZ]
Scope: [systems/regions]
Mitigation: [current action]
IC: [name] | Comms: [name]
Next update: [HH:MM TZ]
Status: [status page link]
SEV1 follow‑up (email or Slack)
Update [N]:
Facts since last: [two bullets max].
Mitigation: [action taken / plan].
ETA confidence: [low/medium/high if safe].
Next update: [HH:MM TZ].
SEV1 resolution
Resolved — [Service/component]
Impact window: [start] to [end, TZ].
Root cause: [known/under investigation].
User impact: [what customers saw].
Remediation: [what changed to resolve].
Follow‑ups: [RCA window and owner].
SEV2 initial
Subject: [SEV2] Degradation on [service]
Impact: [subset of users/functions].
Start: [HH:MM TZ].
Scope: [systems/regions].
Mitigation: [action].
Owner: [IC] | Next update: [HH:MM TZ].
Status page: [link]
Audience variants
Internal engineering
Include ticket links, build IDs, rollback options, and on‑call handoffs. Keep timestamps in UTC.
Executives
Lead with business impact, customer communication, and ETA confidence. One short paragraph is enough.
Customers / status page
Plain language. No stack traces or vendor names. Promise the next update time.
Channel specifics
- Email: Put severity and next update in the subject. Use a short template body.
- Slack: Post in a named incident channel. Pin the latest summary.
- Status page: Log initial, follow‑ups, and resolution. Keep a change log.
Cadence and stop conditions
- SEV1: every 15 minutes until mitigated, then every 30 minutes until resolved.
- SEV2: every 30–60 minutes until resolved.
- Stop when impact ends and the resolution note has been sent.
Make the templates real
- Keep variables consistent:
[service],[impact],[start],[next update],[owner]. - Link the status page. Store these snippets in your on‑call runbook.
- After action: schedule the RCA window and assign owners.
Need help building a practical incident process? Our production engineering team can help you set severity definitions, runbooks, and a clean status‑page flow.