ع
Learn Tracks Reference Guides Saved
playbook

Run customer comms during an incident or outage

Draft the whole incident comms kit in minutes — the status post, the mass reply, the internal single story, the update cadence, the all-clear, and the honest post-incident note — so you spend the outage managing the response, not staring at a blank reply box.

advanced ~prep once, minutes when it matters
when to reach for this

Something just broke for everyone at once, and now support's job is fast, calm, consistent communication across every channel while engineering fixes it. The blank reply box is the enemy — every minute you spend wording the first status post is a minute the inbound flood goes unanswered and every agent improvises a different answer. This playbook drafts the whole comms kit in one sitting so the incident becomes a response you manage rather than a fire you stare at: the initial status post, a templated mass reply, the internal single story so every agent says the exact same thing, a periodic-update cadence, the all-clear, and an honest post-incident note. The discipline that runs through all of it: acknowledge fast, never promise an ETA you can't keep, never speculate on cause or admit fault before the facts are in, and route every external word through whoever owns incident comms.

gather this first
  • The confirmed facts from engineering — what's broken, who and what is affected, and what's being done — and only what's confirmed. Anything still a guess (the cause, the fix time) stays a [placeholder] until it's verified.
  • Who owns sign-off before anything ships externally — usually the incident commander, plus a leadership or legal check on the post-incident note. No customer-facing word goes out without their nod.
  • Your existing voice and macro docs — support-voice.md and macros.md from the support voice and canned-response playbooks — so the incident kit sounds like you under pressure instead of a cold template.
the workflow
  1. Draft the initial status post the moment it's confirmed

    Speed beats polish here, but only on the four things you can actually stand behind: that you see it, who it hits, what you're doing, and when the next update lands. Pin Claude off cause, off blame, and off any ETA you can't keep — those are exactly the lines that come back to bite you.

    you ask
    We have a confirmed incident: [what's broken] is affecting [who/what's affected]. Engineering is [what they're doing]. Draft a short status-page post (under 80 words): acknowledge the problem plainly, state the scope, say we're actively working on it, and commit to a next update by [time]. Do NOT speculate on the cause, do NOT assign blame, and do NOT give a fix-time estimate — use [placeholders] for anything I haven't confirmed.

    what you get back A calm, scoped acknowledgement — "We're aware some customers can't [do X] and are actively working on a fix. Next update by [time]." — with no cause, no blame, and no promised ETA. The post you can publish in two minutes instead of wordsmithing for ten.

    Acknowledging fast is the whole game — a vague-but-honest "we see it, update coming" beats a perfect post that lands twenty minutes late. The technical bug itself is escalated in parallel through the bug-escalation playbook; this step is only the words customers see.

  2. Build the mass-reply macro and a triage rule for the flood

    The inbound spike is mostly the same message fifty times — but not all of it. Get one on-voice mass reply and a rule to separate "affected by this incident" from unrelated tickets, so you reassure the affected without blasting a generic outage notice at someone asking about a refund.

    you ask
    Draft a short mass reply for tickets about this incident, in the voice in support-voice.md: acknowledge, link to the status page for live updates, set the expectation that we'll follow up when it's resolved, and don't over-apologize or admit fault. Then give me a 3-line triage rule — keywords and signals — to tell "affected by this incident" tickets apart from unrelated ones, so we only send this to the right people.

    what you get back A reusable mass-reply macro pointing people to the single source of truth, plus a plain triage rule ("mentions [feature], can't log in, started after [time] → incident; billing/feature questions → normal queue") so the blast lands only on affected tickets.

    Pulling from macros.md keeps the mass reply human; this is the incident extension of your canned-response library, not a cold new template.

  3. Write the internal single story every agent uses

    The fastest way to lose trust mid-incident is five agents giving five different answers. The fix is one internal doc — the agreed facts and the approved language — that every agent reads before they reply, so the customer hears one consistent story no matter who picks up the ticket.

    you ask
    Write an internal incident brief for the support team — not customer-facing. Include: the confirmed facts only, the approved language we're using publicly (and the phrases to avoid, like guessing the cause or promising a fix time), what we can and can't say about [the cause / a timeline], who owns sign-off, and where live updates live. Keep it scannable — agents will read it fast under pressure.

    what you get back A one-screen internal brief: the facts, the do-say / don't-say lines, the boundaries on cause and timeline, and the owner — the single source of truth that makes every agent's reply match every other agent's.

    Mark clearly what's confirmed vs. still unknown. The don't-say list (no cause speculation, no fix-time promises, no fault) is what keeps fifty independent replies from creating fifty small liabilities.

  4. Set the cadence and draft the periodic-update template

    Silence reads as "they've forgotten" even when you're heads-down fixing it. A predictable rhythm — and an update even when there's nothing new — is what holds customer trust through a long incident. Set the interval, then draft the fill-in-the-blank update so each one takes seconds, not a fresh writing session.

    you ask
    Draft a periodic-update template for the status page that I can fill in every [30 minutes]: a one-line current status, what's changed since the last update (or "still investigating" if nothing has), and the time of the next update. Write two example fills — one for "still working, no change" and one for "identified the issue, working on the fix" — both honest, neither promising a fix time.

    what you get back A reusable update template plus two worked examples, including the crucial "still working on it, next update by [time]" one — because a steady drumbeat of honest updates, even empty ones, is what keeps customers from assuming you've gone dark.

  5. Draft the all-clear / resolved message

    When engineering confirms it's fixed — confirmed, not hopeful — you close the loop warmly: it's resolved, here's roughly what happened in plain terms, thanks for the patience, and here's who to contact if you're still seeing trouble. Don't declare victory a minute early; a premature all-clear you have to walk back is worse than waiting.

    you ask
    Engineering has confirmed the incident is resolved. Draft a short all-clear message for the status page and a version for the affected tickets: confirm it's fixed, give a one-sentence plain-language summary of what happened (no jargon, no blame), thank people for their patience, and tell anyone still affected exactly how to reach us. Keep it warm and brief; do not over-promise that it'll never recur.

    what you get back A clean resolved message in both a status-page and a reply flavor — fixed, a human one-line summary, genuine thanks, and a clear path for stragglers — without the over-promise that it can never happen again.

    Only send this after engineering confirms resolution, not when it looks better. Walking back a premature all-clear costs more trust than the extra ten minutes of waiting.

  6. Write the honest post-incident note

    Once the dust settles, the post-incident note is where trust is rebuilt: what happened, the real impact, what you did, and what you're changing so it's less likely next time. Own it plainly — but don't over-apologize into legal risk or promise guarantees you can't keep. This is the one external word that most needs a leadership and legal read before it ships.

    you ask
    Draft a post-incident note for affected customers: what happened (plain language, no deep technical blame), the honest impact and how long it lasted, what we did to fix it, and what we're changing to reduce the chance of a repeat. Own it sincerely without over-apologizing or making any guarantee it can never recur, and flag any line that a legal or leadership reviewer should look at before this goes out.

    what you get back A sincere, specific post-incident note — happened / impact / fixed / changing — that takes responsibility without over-promising, with the legally-sensitive lines flagged for review. The message that turns a bad day into a reason customers trust you more, not less.

    This note carries the most risk of any message in the kit — route it through leadership and legal before it ships. The incident's root cause also feeds your voice-of-customer report and, if it recurs, a help-center article.

make it your own
  • Inherit the calm voice: the de-escalation tone every message in this kit needs comes from Set the support voice and policy every reply inherits — author that once and the incident kit sounds like you under pressure; the mass reply and update templates are the incident extension of Build a canned-response library you'll actually reuse, not cold new copy.
  • Escalate the bug in parallel: this playbook is only the words customers see. The underlying breakage is handed to engineering through Turn a cluster of tickets into a bug engineering will act on at the same time — comms and the fix run on two tracks, not one after the other.
  • Feed the aftermath forward: once it's over, the incident becomes input — it rolls into Turn a month of tickets into a voice-of-customer report as a flagged event, and a pattern that keeps recurring graduates into a Build a help center from the questions you actually get article so customers can self-serve next time. The whole procedure slots into Run a weekly support operating system as the break-glass step you reach for when an outage hits.
  • Pre-stage the kit (Power Track): save the whole flow as a /incident custom command that drafts the status post, mass reply, and internal brief from a one-line "what's broken" input, or a scheduled agent that posts the periodic status update on a timer until you stand it down (see the Playbook's Features tab). Custom commands and scheduled agents are the opt-in Power Track — on Desktop you keep the prompts saved in a note and paste them when the moment hits, which is fast enough that you never need to automate it.
watch out for
  • Never promise an ETA you can't keep. "Next update by [time]" is a promise you control and can always keep; "fixed in 30 minutes" is one engineering hasn't given you — and a missed fix-time estimate destroys more trust than the outage itself. Commit to update cadence, never to a fix time.
  • Never speculate on the cause or admit fault before the facts are in. Early in an incident the cause is a guess, and a guess in writing — "this was a database issue," "our mistake" — becomes a quote that outlives the incident and can carry legal weight. Say what you've confirmed and nothing more; the cause goes public only after engineering verifies it.
  • Coordinate every external word with whoever owns incident comms — eng, leadership, and legal — and route the message through them before it ships. Support doesn't unilaterally narrate an outage: the status post, the all-clear, and especially the post-incident note all clear the incident owner (and legal, for the post-mortem) first. One unreviewed sentence can commit the company to something it can't stand behind.
  • Claude drafts the kit in minutes, but a human owns every word that ships. Speed is the point — and exactly why a person has to confirm each fact, scrub the PII, and approve every external message before it goes live. Claude gives you the calm draft fast; you, not Claude, decide it's true and press publish.

you'll end up with A full incident comms kit drafted in minutes — the initial status post, the mass-reply macro and triage rule, the internal single story, the update-cadence template, the all-clear, and the honest post-incident note — every fact confirmed, every external word routed through the incident owner, so you spend the outage managing a calm, consistent response instead of writing from a blank box while the inbox fills.

Questions people ask

What do I do in the first two minutes of an incident?
Publish the initial status post: acknowledge the problem plainly, state who and what is affected, say you're actively working on it, and commit to a next update by a specific time. Don't wait for a perfect message or for the cause — a fast, honest "we see it, update coming" beats a polished post that lands twenty minutes late, and it stops the inbound flood from going unanswered.
Why shouldn't I give customers a fix-time estimate?
Because engineering rarely knows the real fix time early, and a missed estimate damages trust more than the outage itself. Promise only what you control — the next update time — not the resolution time. "Next update by 3:15" is a promise you can always keep; "fixed by 3:15" is one you can't.
How do I keep every agent from giving a different answer?
Write the internal single story — one brief with the confirmed facts, the approved public language, the phrases to avoid, and who owns sign-off — and have every agent read it before they reply. That single source of truth, plus the mass-reply macro, is what makes fifty replies sound like one consistent voice instead of five conflicting ones.
Does Claude send the incident messages?
No. Claude drafts the whole kit fast so you're not writing from scratch under pressure, but a human confirms every fact, scrubs any customer details, and routes every external word through the incident owner — and leadership and legal for the post-incident note — before it ships. Claude gives you the calm draft; you decide it's true and press publish.