# Incident readiness and runbook coverage

> Whether your playbooks and runbooks cover the incidents you actually have, and what an on-call responder could reach at 3am.

Source: https://triagic.com/checkups/incident-readiness

## Scope [#scope]

Test preparedness against history rather than against a checklist: take the
incidents this org has really had, and ask what a responder would have been able to
find and follow. Local data is enough; an alerting, on-call or source-control
integration makes several steps sharper and the body says where.

## Procedure [#procedure]

1. **Recover the incident history.** Using the ticket data available to you, find
   the incident-shaped events of the last 90–180 days: multiple tickets about the
   same failure in a short window, tickets whose investigations reached an
   infrastructure or platform root cause, anything marked or escalated as urgent.
   `tickets__search_similar` is the fastest route to the clusters. Where an
   alerting or on-call integration is connected, pull its incident list too and
   reconcile the two — incidents that exist in the alerting system with no
   corresponding tickets, and ticket storms with no alert, are both findings before
   any coverage question is asked. Produce a plain list: what happened, when, how
   long, what the root cause turned out to be.
2. **Inventory the response material.** List the playbooks (name, instructions,
   routing hints, scoped servers, last updated) and any runbook content the org
   keeps where this run can see it. For each, note what class of incident it is
   written for and how recently it was touched.
3. **Cross-tabulate history against material — this is the checkup.** For each
   incident from step 1, answer three questions:
   * Did material exist that covered it?
   * Would it have been *found*? A playbook whose routing hints do not contain the
     words the incident's tickets actually used will not be reached under pressure,
     even though it exists.
   * Would it have *worked*? Compare the steps in the material against what the
     investigation actually did. Material naming a system that is no longer
     connected, or an owner who has left, or a dashboard that no longer exists, is
     worse than none: it costs a responder time before it fails them.
     The answer "covered, findable, and it matches what we actually did" is the only
     pass. Anything else is a gap, and say which of the three it failed on.
4. **The cold-start test.** Take the two or three most likely incident types and
   walk them as a responder with no context would: from the alert or first ticket,
   what is the first thing to look at, and is it reachable through the connected
   integrations right now? Note every step that requires knowledge held by one
   person, access nobody on call has, or a system Triagic cannot see. Single points
   of human failure are the most common serious finding in this checkup and they
   never appear in a coverage table.
5. **Detection and escalation surface.** Ask how each historical incident was
   first noticed: an alert, or a customer ticket. Incidents first reported by
   customers are a detection gap regardless of how well the response then went, and
   the ratio of the two over the window is the single best readiness number this
   checkup can produce. Where an on-call integration is connected, check that a
   schedule exists, that it has no gaps, and that escalation goes somewhere other
   than one person; where none is connected, ask how the team is reached out of
   hours and record the answer as not assessable if nothing shows it.
6. **Post-incident follow-through.** For each incident, look for what was recorded
   afterwards: a root cause written down, a fix, a note added to a playbook. An
   incident with no durable artifact will be re-solved from scratch when it recurs,
   and a recurrence you can already see in step 1 proves it.
7. **Freshness.** Flag playbooks not touched in six months that cover systems which
   have changed since, and any material whose scoped servers no longer exist in the
   integration inventory.

## Finding keys [#finding-keys]

The `key` names the *coverage or readiness gap*, not the incident and not this
run, so a gap that is still open next month keeps one ledger row and closes when
the material is written. Where the gap is about a specific playbook, key on the
playbook.

* `coverage:database-failover:no-playbook`
* `coverage:payment-gateway-outage:playbook-not-findable`
* `playbook:cache-incidents:steps-reference-dead-system`
* `detection:customer-first-reported:majority`
* `oncall:escalation:single-person`
* `followup:incident-2026-07-14:no-root-cause-recorded`

## Severity rubric [#severity-rubric]

Severity is about what it would cost during the next incident.

* **critical** — a recurring incident type with no material at all, or material
  whose steps would actively mislead a responder (dead systems, departed owners);
  escalation that depends on one unavailable person.
* **high** — most incidents first reported by customers rather than detected; a
  playbook that exists but would not be found from the words real tickets use; no
  durable record after a serious incident that has already recurred.
* **medium** — coverage that exists but is stale, incomplete, or scoped to systems
  that have moved; a cold-start step that depends on knowledge one person holds.
* **low** — freshness and tidiness: material not reviewed in a long time but still
  correct; naming that obscures which incident type a playbook covers.
* **info** — the coverage matrix itself and the detection ratio, recorded as the
  readiness baseline for the next run.

## Output guidance [#output-guidance]

Open with a one-paragraph executive summary: how many incidents in
the window, how many would have had usable material, the detection ratio (alert-
first versus customer-first), and the largest gap. Then: an incident history table
(Incident | When | Duration | Root cause | First detected by), the coverage matrix
(Incident type | Material exists? | Would be found? | Would it work? | Verdict) as
the centrepiece, a Cold-start section walking the top incident types step by step
and naming every point where a responder would stall, and a Detection and
escalation section. Close with a recommendations table (Gap | Recommended action |
Owner | Effort), ordered by what it would cost during the next incident rather than
by effort. Where something could not be assessed because nothing is connected, say
so plainly instead of assuming the org is unprepared.

<!-- generated by apps/server/scripts/export-checkups.ts, do not edit -->
