# SLA and response quality

> Breached targets, slow first responses, tickets left to rot, and reply-quality outliers, from your own ticket history.

Source: https://triagic.com/checkups/sla-response-quality

## Scope [#scope]

A weekly read of how the support operation is actually performing, entirely from
Triagic's own ticket data — no integration required. Measure response and
resolution timing, find the tickets that fell through, and sample reply quality.
Judge the *system*, not individuals: the finding is "tickets arriving through
channel X wait twice as long", not "this person is slow".

## Procedure [#procedure]

1. **Fix the window and the targets, and say both.** Default to the last 7 days of
   tickets, with the previous 7 as the comparison. Then state the response targets
   you are measuring against. If the org has recorded targets, use them. If it has
   not — most have not — say so explicitly and use these as the stated working
   assumption: first response within 4 business hours, resolution within 2 business
   days for ordinary tickets and 4 hours for anything incident-shaped. Every
   "breach" in the report must be traceable to a target the report itself printed.
2. **Build the timeline per ticket.** Using the ticket data available to you, take
   each ticket's creation time, its current status, its last update, and its
   activity timeline (notes, emails, events) where one is present. Derive three
   numbers per ticket: time to first outbound response (the first reply that went
   to the requester, not an internal note), time to resolution for anything now
   resolved, and total age for anything still open. Where the activity timeline is
   missing or thin, say which tickets could not be measured and exclude them from
   the aggregates rather than treating "no recorded reply" as an infinite wait.
3. **Distributions, not averages.** Report median and p90 for first response and
   for resolution — a mean hides exactly the tail this checkup exists to find. Then
   break the same numbers down by ticket source (`hubspot`, `zendesk`, `jira`,
   `slack`, `thread`, `manual`) and, where the data supports it, by playbook.
   A channel with a much worse p90 than the others is usually a routing or
   notification gap, not an effort gap — check whether that source's tickets are
   even reaching the same queue.
4. **The tail is the finding.** List every ticket that breached, worst first, with
   its age and current status. Then look for the shapes that matter more than the
   individual breaches:
   * **Rot** — tickets in `new` for longer than the first-response target: nobody
     has picked them up at all. This is the most actionable list in the report.
   * **Stalls** — tickets in `investigating` with no activity for days: work
     started and stopped.
   * **Silent failures** — tickets with status `failed`: automated triage gave up
     and, unless someone noticed, nothing replaced it.
   * **Reopens** — a ticket that returned to an open status after being resolved
     means the first resolution did not hold, and reopen rate is a better quality
     signal than resolution time.
5. **Sample quality, do not score everyone.** Take a small sample — say ten —
   spread across fast and slow tickets and read the actual replies. Assess against
   things a reader can verify: did the first reply answer the question asked or
   only acknowledge it; did it ask for information the customer had already given;
   was the customer told what happens next and by when; did the tone match the
   customer's state (a frustrated second-time reporter given a template reply is a
   finding); did an internal detail leak into a customer-facing message. Quote at
   most a short excerpt, and never customer personal data. Report patterns across
   the sample; a single awkward sentence is not a finding.
6. **Load context before concluding.** Compare volume this window against the
   last: a p90 that doubled in a week where volume also doubled is a capacity
   story, and recommending "respond faster" would be the wrong answer. Say which
   of the two it is.
7. **Close the loop with the previous run.** Breaches that recur — same source,
   same shape, week after week — are a process finding and outrank a worse-looking
   one-off spike.

## Finding keys [#finding-keys]

The `key` names the *recurring problem*, not the individual ticket or this week,
so a channel that is still slow next week updates one ledger row and a fixed
channel that slips again is caught as a regression. Key on the pattern; put the
week's specific ticket ids in `evidence`, which is refreshed each run. A single
egregious ticket may have its own key when it genuinely is the finding.

* `sla:first-response:p90-breach`
* `source:zendesk:first-response-slow`
* `queue:new:rot-over-target`
* `status:failed:unrecovered-triage`
* `quality:first-reply:no-next-step`
* `ticket:4821:stalled-14d`

## Severity rubric [#severity-rubric]

Severity is about customer harm and the risk of it repeating, not about the
absolute number of hours.

* **critical** — tickets abandoned entirely: anything sitting in `new` or
  `failed` past several multiples of the target with no human touch, or an
  incident-shaped ticket that got an ordinary-ticket response time.
* **high** — a systematic breach rather than an outlier: p90 first response beyond
  target for a whole source or the whole window, or a reopen rate that says
  resolutions are not holding.
* **medium** — a meaningful minority of tickets breaching, a stalled cohort in
  `investigating`, or a repeated quality pattern in the sample (replies that
  acknowledge without answering, missing next steps).
* **low** — isolated breaches with a visible reason, small tone or clarity issues,
  metrics slightly worse than last week within normal variance.
* **info** — the week's numbers themselves, recorded as the trend baseline, and any
  measurable improvement worth naming.

## Output guidance [#output-guidance]

Open with a one-paragraph executive summary in plain words: volume
this window versus last, median and p90 first response, how many tickets breached,
and the single worst pattern. Print the targets being measured against immediately
after, and say explicitly whether they are the org's own or this checkup's stated
assumption. Then sections for Timing, The tail (rot, stalls, failed, reopens),
By source and playbook, and Reply quality — the tail section carries the actual
ticket list, worst first. Close with a recommendations table (Problem | Evidence |
Recommended change | Owner | Effort), process changes above exhortations to be
faster. Never rank or name individual agents; every finding is about a channel, a
queue, a stage, or a pattern. Say "the week met its targets" outright when it did.

<!-- generated by apps/server/scripts/export-checkups.ts, do not edit -->
