# Ticket trends and recurring root causes

> What volume is doing, the anatomy of each spike, and the root causes that keep coming back across months.

Source: https://triagic.com/checkups/ticket-trend-rca

## Scope [#scope]

Read the ticket history as a signal about the product, not as a workload. Two
questions: what happened to volume and why, and which root causes recur often
enough that fixing the cause would be cheaper than handling the tickets. Local
data only; richer where a product-analytics or error-tracking integration is
connected, but complete without one.

## Procedure [#procedure]

1. **Establish the shape of volume.** Using the ticket data available to you, count
   tickets per day (or per week for a long window) over the last 90 days, and
   compare the most recent 30 against the preceding 30. Break the same series down
   by source (`hubspot`, `zendesk`, `jira`, `slack`, `thread`, `manual`).
   Say what normal looks like before saying anything is abnormal: give the typical
   daily range and the weekday/weekend pattern, so a "spike" is measured against a
   baseline the report actually printed.
2. **Find the spikes.** A day materially above the recent typical range is a
   candidate. Ignore ones fully explained by a change in a single source's ingest
   (a connector that was down for two days and then backfilled produces a fake
   spike — check for that before investigating a cause).
3. **Anatomy of each real spike, in this order.** For each:
   a. **Composition** — what were the tickets about? Cluster the spike day's
   tickets by root cause and subject; a spike is almost always one dominant
   cluster plus normal background. Name the dominant cluster and its share.
   b. **Onset** — when did the first ticket of that cluster arrive, to the hour?
   That timestamp, not the peak, is what any candidate cause has to precede.
   c. **Candidate cause** — look for something that changed just before the onset.
   Where source control or CI is connected, check deploys and merges in that
   window; where error tracking is connected, check for a new issue or a rate
   change starting at the same hour; where product analytics is connected, check
   whether a funnel step's conversion dropped at the same time. With none of
   those connected, say the correlation could not be tested and give the onset
   timestamp so a human can check it in minutes.
   d. **Resolution** — how did the cluster end: a fix, a workaround, or did the
   tickets simply stop arriving with no recorded resolution? The third case is
   the most important to flag, because the cause is still out there.
4. **Recurring root causes across the whole window.** Independently of spikes,
   group the investigated root causes over 90 days and rank by frequency and by
   total investigation effort. `tickets__search_similar` is the way to find the
   recurrences that are worded differently each time. For each recurring cause,
   record: how many tickets, over what span, whether it is accelerating, and
   whether the fix has ever been described as permanent in any investigation. A
   root cause that recurs monthly and is "fixed" every time is a signal that the
   real cause was never found — say that.
5. **Chronic versus acute.** Separate the two, because they need different owners.
   Acute is a spike with a cause and an end. Chronic is a steady trickle that never
   spikes and never stops — chronic causes are systematically under-noticed
   precisely because they never create a bad day, and adding up their annual ticket
   count is usually the most persuasive number in this report.
6. **Look for what stopped.** A cluster that used to arrive weekly and has not
   appeared in a month is worth recording too, both as a genuine win and as a
   prompt to check that the tickets are not simply being lost before they reach the
   queue.
7. **Cost the top causes.** For each recurring cause, multiply ticket count by the
   typical resolution time to give a rough handling cost over the window. Say
   plainly that it is an estimate; it exists to rank causes against each other, not
   to be quoted as a budget figure.

## Finding keys [#finding-keys]

The `key` names the *cause or pattern*, not a date or a spike, so a chronic cause
keeps one ledger row across months and closes when it genuinely stops. Give each
recurring cause a stable slug and reuse it; a spike gets a key only when it turned
out to have an unresolved cause worth tracking.

* `cause:payout-bank-verification:recurring`
* `cause:oauth-token-expiry:chronic`
* `spike:checkout-declines:unresolved-cause`
* `trend:volume:rising-30d`
* `source:slack:ingest-gap`

## Severity rubric [#severity-rubric]

Severity is about the cause's ongoing cost and the risk it repeats, not about how
loud the spike was.

* **critical** — a spike whose cause was never identified and whose tickets simply
  stopped arriving; or a recurring cause that is accelerating and customer-visible.
* **high** — a chronic cause among the top few by handling cost that no fix has
  ever addressed; a sustained rise in volume with no explanation.
* **medium** — a recurring cause with a known workaround but no permanent fix; a
  spike that was explained and resolved but could recur because nothing changed.
* **low** — small recurring annoyances, seasonal or expected variation worth
  recording.
* **info** — the baseline itself, resolved trends, and clusters that stopped —
  recorded so the next run can compare.

## Output guidance [#output-guidance]

Open with a one-paragraph executive summary in plain words: volume
this period versus last, the number of spikes and whether their causes were found,
and the single recurring cause costing the most. Then: a Volume section that prints
the baseline before any anomaly claim, a section per spike following the
composition / onset / candidate cause / resolution structure with the onset
timestamp stated exactly, a Recurring root causes table (Cause | Tickets | Span |
Accelerating? | Estimated handling cost | Ever permanently fixed?), and a short
What stopped section. Close with a recommendations table (Cause | Recommended fix |
Who owns it | Tickets avoided per month). Distinguish everywhere between what the
data shows and what you inferred, and where a correlation could not be tested
because nothing is connected, say so and give the timestamp a human needs.

<!-- generated by apps/server/scripts/export-checkups.ts, do not edit -->
