Skip to content

Knowledge and playbook gaps

Recurring ticket clusters that no playbook covers, plus drafted playbook instructions for the largest gaps.

Scope

Find the work the team keeps redoing from scratch. Cluster the recent ticket history by what actually went wrong, compare those clusters against the playbooks that exist, and draft the instructions for the biggest uncovered cluster. Local data only; no integration needed.

Procedure

  1. Read the existing coverage first. Using the playbook data available to you, list every playbook: name, what its routing hints say it catches, the systems it scopes to, and when it was last updated. This is the map the rest of the run compares against, and reading it first stops you "discovering" a gap that a playbook already covers under a different name.
  2. Cluster the recent tickets by root cause, not by subject. Work over the last 30–90 days of ticket history. Subject lines cluster badly — the same failure arrives as "can't log in", "app broken", and "urgent!!" — so cluster on what the investigation concluded: the recorded root cause where one exists, and the body plus resolution where it does not. tickets__search_similar is the tool for this: take a representative ticket, pull its neighbours, and let the cluster grow from there; repeat from tickets that no cluster has claimed yet. Stop when the remaining tickets are genuine one-offs. For each cluster record: a plain name, size, the span of dates, whether it is growing, and two or three representative ticket ids.
  3. Match clusters to playbooks. For every cluster, decide which of three it is:
    • Covered — a playbook exists and the tickets in the cluster were actually routed to it. Nothing to do.
    • Covered but not routed — a playbook exists that would have handled these tickets, and they went elsewhere or nowhere. This is a routing finding, not a knowledge gap, and the fix is the playbook's routing hints, not new content. It is also the cheapest fix in this checkup, so look for it deliberately.
    • Uncovered — no playbook addresses it. This is the gap.
  4. Rank the gaps by cost, not by size. A cluster's real cost is roughly its volume multiplied by how long its tickets take to resolve, weighted up if it is growing and up again if its tickets escalate or reopen. A ten-ticket cluster that each time takes two days of investigation outranks a fifty-ticket cluster resolved by a one-line reply.
  5. Check whether the knowledge exists but is trapped. Before declaring a gap, look at how the cluster's tickets were actually resolved: if several investigations independently reached the same root cause, the knowledge exists — it is just re-derived every time and lives nowhere. Say that explicitly. A trapped-knowledge gap is easier to close than a real unknown, because the draft in the next step can be written straight from the resolved tickets.
  6. Draft, do not just recommend. For the top two or three gaps, write an actual playbook draft in the report: a name, routing hints (the words and phrases these tickets really use, taken from the cluster, not invented), the numbered investigation steps a competent agent should follow — naming the systems and the order to check them, drawn from what the successful investigations in the cluster actually did — and the set of servers the playbook should be scoped to. Someone should be able to paste it into the playbook editor and have it work.
  7. Check the other direction too. Playbooks that caught nothing in the window are also a finding: either their routing hints no longer match how tickets are worded, or the problem they cover has gone away. Say which you think it is and why.

Finding keys

The key names the knowledge gap, not this week's tickets, so a gap that is still open next week updates one row — and closes automatically once a playbook covers it. Name the cluster with a stable slug you will use again; keep the ticket ids in evidence, which is refreshed every run.

  • gap:payout-bank-verification:no-playbook
  • gap:oauth-token-expiry:no-playbook
  • routing:checkout-declines:hints-miss-cluster
  • trapped:webhook-retry-exhaustion:knowledge-not-written-down
  • playbook:legacy-imports:unused-90d

Severity rubric

Severity is the ongoing cost of the gap, not the difficulty of closing it.

  • critical — a large, growing cluster with no coverage whose tickets are incident-shaped or customer-escalating: the team is re-improvising a response to something that keeps happening.
  • high — a top cluster by cost with no playbook, or a cluster whose knowledge is demonstrably trapped (several independent investigations reaching the same root cause).
  • medium — a real but smaller uncovered cluster; a routing miss where a good playbook exists and is not being reached.
  • low — a small cluster worth a note in an existing playbook rather than a new one; a playbook whose hints are drifting.
  • info — coverage summary: clusters found, share of tickets covered, and playbooks that matched nothing this window.

Output guidance

Open with a one-paragraph executive summary: how many tickets were clustered, how many clusters, what share of the volume is covered by an existing playbook, and the single largest uncovered cluster. Then a clusters table (Cluster | Tickets | Trend | Coverage | Example ticket ids), ranked by cost rather than size. Then a section per top gap containing a complete, paste-ready playbook draft — name, routing hints, numbered instructions, suggested server scope. Then a short section on playbooks that matched nothing. Close with a recommendations table (Action | Cluster | Expected effect | Effort). Everything must be traceable to real tickets: cite ids, and never invent a root cause the investigations do not support.

Run it against your systems

This checkup is in the desktop app under Checkups. No card, read-only credentials you configure.