How to write incident updates customers trust (with templates)

During an outage, what you say matters almost as much as how fast you fix it. A simple four-stage structure, ready-to-use templates, and the mistakes that make customers lose trust.

· 4 min read · By the Spot Downtime team

Customers forgive outages. What they don't forgive is silence, vague reassurance, or finding out from Twitter before they hear it from you. A clear, honest stream of updates during an incident can leave customers trusting you more than before it started.

You don't need a communications team for this. You need a simple structure, a few templates prepared in advance, and the discipline to post the first update quickly.

The four stages of an incident

Almost every incident moves through the same stages. Naming them tells readers where you are at a glance:

  • Investigating: you know something is wrong, not yet why.
  • Identified: you know the cause and are working on a fix.
  • Monitoring: a fix is in place; you're watching to be sure it holds.
  • Resolved: everything is back to normal.

Not every incident needs all four. A five-minute blip can go straight from Investigating to Resolved.

The first update: fast beats complete

The first update has one job: show you know. Post it within 5 to 10 minutes of noticing the problem, even if you know almost nothing yet. Every minute of silence turns into support tickets and guesswork.

A good first update answers three questions:

  • What's affected, in customer terms (“signing in”, not “the auth service”)?
  • What are you doing?
  • When is the next update?

Templates

Copy these into your runbook and fill in the brackets. Plain language, no jargon, no blame.

Investigating

text
We're looking into reports that [some customers can't sign in / the dashboard is loading slowly].
[Other features / the API] are working normally. We'll post another update within 30 minutes.

Identified

text
We've found the cause: [a database server ran out of disk space]. We're [freeing up space and
restarting the affected service]. Sign-ins are still failing for some customers.
Next update within 30 minutes, or sooner if it's fixed.

Monitoring

text
A fix is in place and sign-ins are working again. We're keeping a close eye on things
to make sure it holds. If you're still having trouble, please try again or contact support.

Resolved

text
This incident is resolved. Between [10:42] and [11:27 UTC], some customers couldn't sign in
because [a database server ran out of disk space]. No data was lost. We've [added alerts on
disk usage] so this can't happen silently again. We're sorry for the disruption.

Scheduled maintenance

text
On [Tuesday 14 October, 02:00–02:30 UTC] we're upgrading our database. [The dashboard] may be
unavailable for a few minutes during that window. [Monitoring and alerts] are not affected.

Write for the reader, not the engineer

  • Describe impact, not internals. “Some customers can't receive alert emails” helps; “SMTP relay degradation in eu-west-1” doesn't.
  • Say who is affected. “Customers in Europe” or “about 10% of sign-ins” is much better than “some users”, if you know it.
  • Use times with a time zone. “Since 10:42 UTC”, never “since this morning”.
  • Promise update times, not fix times. “Next update in 30 minutes” is a promise you can keep. “Fixed within the hour” often isn't.
  • Keep every promise. If the next update is due and nothing changed, say so: “Still working on it, no change yet, next update in 30 minutes.” Silence after a promised update is worse than no promise at all.

Mistakes that cost trust

  • “All systems operational” during an outage. A status page that disagrees with what customers see is worse than none. Automate it: let monitoring mark services as down.
  • Minimising. “A small number of users may be experiencing intermittent issues”, when everyone is down, reads as evasive.
  • Blaming others. Even if your cloud provider caused it, your customers chose you. “Our hosting provider had an outage, and we're adding redundancy so one provider can't take us down” is fine. “It's AWS's fault” isn't.
  • Going quiet after the fix. The resolved update is the one people remember. Explain what happened and what you're changing.

Keep two channels

Your team needs the technical details; your customers need the impact. Post internal notes for the team and public updates for customers, so neither has to read the other's version.

Put it where people will look

Updates are only useful if customers can find them. A public status page at a predictable address (like status.yourcompany.com) becomes the first place people check, and email subscriptions bring updates to the people who want them without anyone refreshing a page.

In Spot Downtime, this is built into every incident:

  • When a monitor fails, an incident opens automatically and your status page shows the affected service as down.
  • Your team posts updates with the stages above. Each update is either public (shown on the status page and emailed to subscribers) or an internal note for the team only.
  • Updates also go to your alert channels, so everyone working on the incident sees the latest.
  • Scheduled maintenance is announced on the status page in advance, and checks pause during the window.

See status pages or create one for free. It takes about two minutes, and your first status page is included in every plan.

Keep reading