Incident readiness and runbook coverage
Whether your playbooks and runbooks cover the incidents you actually have, and what an on-call responder could reach at 3am.
Scope
Test preparedness against history rather than against a checklist: take the incidents this org has really had, and ask what a responder would have been able to find and follow. Local data is enough; an alerting, on-call or source-control integration makes several steps sharper and the body says where.
Procedure
- Recover the incident history. Using the ticket data available to you, find
the incident-shaped events of the last 90–180 days: multiple tickets about the
same failure in a short window, tickets whose investigations reached an
infrastructure or platform root cause, anything marked or escalated as urgent.
tickets__search_similaris the fastest route to the clusters. Where an alerting or on-call integration is connected, pull its incident list too and reconcile the two — incidents that exist in the alerting system with no corresponding tickets, and ticket storms with no alert, are both findings before any coverage question is asked. Produce a plain list: what happened, when, how long, what the root cause turned out to be. - Inventory the response material. List the playbooks (name, instructions, routing hints, scoped servers, last updated) and any runbook content the org keeps where this run can see it. For each, note what class of incident it is written for and how recently it was touched.
- Cross-tabulate history against material — this is the checkup. For each
incident from step 1, answer three questions:
- Did material exist that covered it?
- Would it have been found? A playbook whose routing hints do not contain the words the incident's tickets actually used will not be reached under pressure, even though it exists.
- Would it have worked? Compare the steps in the material against what the investigation actually did. Material naming a system that is no longer connected, or an owner who has left, or a dashboard that no longer exists, is worse than none: it costs a responder time before it fails them. The answer "covered, findable, and it matches what we actually did" is the only pass. Anything else is a gap, and say which of the three it failed on.
- The cold-start test. Take the two or three most likely incident types and walk them as a responder with no context would: from the alert or first ticket, what is the first thing to look at, and is it reachable through the connected integrations right now? Note every step that requires knowledge held by one person, access nobody on call has, or a system Triagic cannot see. Single points of human failure are the most common serious finding in this checkup and they never appear in a coverage table.
- Detection and escalation surface. Ask how each historical incident was first noticed: an alert, or a customer ticket. Incidents first reported by customers are a detection gap regardless of how well the response then went, and the ratio of the two over the window is the single best readiness number this checkup can produce. Where an on-call integration is connected, check that a schedule exists, that it has no gaps, and that escalation goes somewhere other than one person; where none is connected, ask how the team is reached out of hours and record the answer as not assessable if nothing shows it.
- Post-incident follow-through. For each incident, look for what was recorded afterwards: a root cause written down, a fix, a note added to a playbook. An incident with no durable artifact will be re-solved from scratch when it recurs, and a recurrence you can already see in step 1 proves it.
- Freshness. Flag playbooks not touched in six months that cover systems which have changed since, and any material whose scoped servers no longer exist in the integration inventory.
Finding keys
The key names the coverage or readiness gap, not the incident and not this
run, so a gap that is still open next month keeps one ledger row and closes when
the material is written. Where the gap is about a specific playbook, key on the
playbook.
coverage:database-failover:no-playbookcoverage:payment-gateway-outage:playbook-not-findableplaybook:cache-incidents:steps-reference-dead-systemdetection:customer-first-reported:majorityoncall:escalation:single-personfollowup:incident-2026-07-14:no-root-cause-recorded
Severity rubric
Severity is about what it would cost during the next incident.
- critical — a recurring incident type with no material at all, or material whose steps would actively mislead a responder (dead systems, departed owners); escalation that depends on one unavailable person.
- high — most incidents first reported by customers rather than detected; a playbook that exists but would not be found from the words real tickets use; no durable record after a serious incident that has already recurred.
- medium — coverage that exists but is stale, incomplete, or scoped to systems that have moved; a cold-start step that depends on knowledge one person holds.
- low — freshness and tidiness: material not reviewed in a long time but still correct; naming that obscures which incident type a playbook covers.
- info — the coverage matrix itself and the detection ratio, recorded as the readiness baseline for the next run.
Output guidance
Open with a one-paragraph executive summary: how many incidents in the window, how many would have had usable material, the detection ratio (alert- first versus customer-first), and the largest gap. Then: an incident history table (Incident | When | Duration | Root cause | First detected by), the coverage matrix (Incident type | Material exists? | Would be found? | Would it work? | Verdict) as the centrepiece, a Cold-start section walking the top incident types step by step and naming every point where a responder would stall, and a Detection and escalation section. Close with a recommendations table (Gap | Recommended action | Owner | Effort), ordered by what it would cost during the next incident rather than by effort. Where something could not be assessed because nothing is connected, say so plainly instead of assuming the org is unprepared.
Run it against your systems
This checkup is in the desktop app under Checkups. No card, read-only credentials you configure.