# Knowledge and playbook gaps

> Recurring ticket clusters that no playbook covers, plus drafted playbook instructions for the largest gaps.

Source: https://triagic.com/checkups/knowledge-gaps

## Scope [#scope]

Find the work the team keeps redoing from scratch. Cluster the recent ticket
history by what actually went wrong, compare those clusters against the playbooks
that exist, and draft the instructions for the biggest uncovered cluster. Local
data only; no integration needed.

## Procedure [#procedure]

1. **Read the existing coverage first.** Using the playbook data available to you,
   list every playbook: name, what its routing hints say it catches, the systems it
   scopes to, and when it was last updated. This is the map the rest of the run
   compares against, and reading it first stops you "discovering" a gap that a
   playbook already covers under a different name.
2. **Cluster the recent tickets by root cause, not by subject.** Work over the last
   30–90 days of ticket history. Subject lines cluster badly — the same failure
   arrives as "can't log in", "app broken", and "urgent!!" — so cluster on what the
   investigation concluded: the recorded root cause where one exists, and the body
   plus resolution where it does not. `tickets__search_similar` is the tool for
   this: take a representative ticket, pull its neighbours, and let the cluster
   grow from there; repeat from tickets that no cluster has claimed yet. Stop when
   the remaining tickets are genuine one-offs. For each cluster record: a plain
   name, size, the span of dates, whether it is growing, and two or three
   representative ticket ids.
3. **Match clusters to playbooks.** For every cluster, decide which of three it is:
   * **Covered** — a playbook exists and the tickets in the cluster were actually
     routed to it. Nothing to do.
   * **Covered but not routed** — a playbook exists that would have handled these
     tickets, and they went elsewhere or nowhere. This is a *routing* finding, not
     a knowledge gap, and the fix is the playbook's routing hints, not new content.
     It is also the cheapest fix in this checkup, so look for it deliberately.
   * **Uncovered** — no playbook addresses it. This is the gap.
4. **Rank the gaps by cost, not by size.** A cluster's real cost is roughly its
   volume multiplied by how long its tickets take to resolve, weighted up if it is
   growing and up again if its tickets escalate or reopen. A ten-ticket cluster
   that each time takes two days of investigation outranks a fifty-ticket cluster
   resolved by a one-line reply.
5. **Check whether the knowledge exists but is trapped.** Before declaring a gap,
   look at how the cluster's tickets were actually resolved: if several
   investigations independently reached the same root cause, the knowledge exists —
   it is just re-derived every time and lives nowhere. Say that explicitly. A
   trapped-knowledge gap is easier to close than a real unknown, because the draft
   in the next step can be written straight from the resolved tickets.
6. **Draft, do not just recommend.** For the top two or three gaps, write an actual
   playbook draft in the report: a name, routing hints (the words and phrases these
   tickets really use, taken from the cluster, not invented), the numbered
   investigation steps a competent agent should follow — naming the systems and
   the order to check them, drawn from what the successful investigations in the
   cluster actually did — and the set of servers the playbook should be scoped to.
   Someone should be able to paste it into the playbook editor and have it work.
7. **Check the other direction too.** Playbooks that caught nothing in the window
   are also a finding: either their routing hints no longer match how tickets are
   worded, or the problem they cover has gone away. Say which you think it is and
   why.

## Finding keys [#finding-keys]

The `key` names the *knowledge gap*, not this week's tickets, so a gap that is
still open next week updates one row — and closes automatically once a playbook
covers it. Name the cluster with a stable slug you will use again; keep the ticket
ids in `evidence`, which is refreshed every run.

* `gap:payout-bank-verification:no-playbook`
* `gap:oauth-token-expiry:no-playbook`
* `routing:checkout-declines:hints-miss-cluster`
* `trapped:webhook-retry-exhaustion:knowledge-not-written-down`
* `playbook:legacy-imports:unused-90d`

## Severity rubric [#severity-rubric]

Severity is the ongoing cost of the gap, not the difficulty of closing it.

* **critical** — a large, growing cluster with no coverage whose tickets are
  incident-shaped or customer-escalating: the team is re-improvising a response to
  something that keeps happening.
* **high** — a top cluster by cost with no playbook, or a cluster whose knowledge is
  demonstrably trapped (several independent investigations reaching the same root
  cause).
* **medium** — a real but smaller uncovered cluster; a routing miss where a good
  playbook exists and is not being reached.
* **low** — a small cluster worth a note in an existing playbook rather than a new
  one; a playbook whose hints are drifting.
* **info** — coverage summary: clusters found, share of tickets covered, and
  playbooks that matched nothing this window.

## Output guidance [#output-guidance]

Open with a one-paragraph executive summary: how many tickets were
clustered, how many clusters, what share of the volume is covered by an existing
playbook, and the single largest uncovered cluster. Then a clusters table (Cluster |
Tickets | Trend | Coverage | Example ticket ids), ranked by cost rather than size.
Then a section per top gap containing a complete, paste-ready playbook draft —
name, routing hints, numbered instructions, suggested server scope. Then a short
section on playbooks that matched nothing. Close with a recommendations table
(Action | Cluster | Expected effect | Effort). Everything must be traceable to real
tickets: cite ids, and never invent a root cause the investigations do not support.

<!-- generated by apps/server/scripts/export-checkups.ts, do not edit -->
