# Webhook and integration failure deep-dive

> Run this against a named ticket or incident: trace a webhook or integration failure end to end and find where events are being lost.

Source: https://triagic.com/checkups/webhook-failure-triage

## Scope [#scope]

A deep-dive triage procedure for events that were sent and never took effect:
webhook deliveries, queue consumers, sync jobs, integration callbacks. Run on
demand with a **focus** naming what to investigate. Read-only throughout — never
replay, re-drive, re-queue, acknowledge, or purge anything, however tempting; say
what should be replayed and let a human do it.

## Procedure [#procedure]

1. **Start from the focus.** The operator names a ticket or an incident — "ticket
   \#4821", "the Stripe webhook backlog from Tuesday", "orders stuck since the
   deploy". Begin by loading it: pull that ticket, its body, its metadata and its
   investigation history, and restate in one line what is claimed to be broken,
   for whom, and since when. Everything after this is anchored to those three
   facts. Use `tickets__search_similar` on the focus text immediately — this
   exact failure has often happened before, and a resolved neighbour usually
   contains the answer and shortens the whole run.
   **If no focus was given**, do not stop. Fall back to finding the subject
   yourself: look for the most recent cluster of integration- or webhook-shaped
   tickets, or the most recent failed sync visible in the data available to you.
   State clearly at the top of the report which incident you selected and why, and
   invite the operator to re-run with an explicit focus if you picked wrong.
2. **Establish the boundary conditions before tracing.** Pin down the exact time
   window (first failure to now, or to the last known good event), the direction
   (inbound webhook the org receives, or outbound call the org makes), the volume
   (one event or thousands), and the blast radius (one customer or everyone). A
   single-customer failure and a total outage look identical in a ticket and need
   completely different investigations — get this wrong and the rest is wasted.
3. **Walk the delivery chain in order and find the last hop that has the event.**
   The chain is: sender → transport/edge → receiver endpoint → queue → consumer →
   application state. Confirm the event's presence or absence at each hop before
   moving to the next, and stop where it disappears — that hop is the fault, and
   everything downstream is a symptom. Concretely, with whatever is connected:
   * **Sender** — the provider's own delivery log or event list, with the response
     code it recorded. A provider showing 200s has done its job; a provider showing
     4xx/5xx or retry exhaustion has just told you the answer.
   * **Transport and edge** — CDN, proxy or gateway logs for the endpoint path in
     the window: TLS failures, 403s from a rule change, body-size or timeout
     limits. Signature-verification rejections live here or at the receiver and are
     a very common cause after a secret rotation.
   * **Receiver** — application logs and error tracking around each failure
     timestamp. Look for the handler's own errors, and for a response that returned
     2xx *before* durably persisting the event, which loses events silently and is
     the nastiest shape in this whole class.
   * **Queue** — depth, age of the oldest message, consumer lag, dead-letter depth
     and in-flight counts. A growing dead-letter queue is the loudest possible
     evidence and is routinely never looked at.
   * **Consumer** — is it running, is it erroring, is it crash-looping, was it
     scaled to zero, is it stuck on one poison message blocking a partition?
   * **Application state** — query for the records the events should have created
     or updated and confirm the gap the customer is reporting is real. Do this last
     but always do it: it converts a plausible story into a verified one.
4. **Anchor to a change.** Almost every failure of this class starts at a change.
   Ask what changed just before the first failure: a deploy, a config or secret
   rotation, a certificate expiry, a provider API version bump, an IP allowlist or
   firewall edit, a scaling change. Where source control or CI is connected, check
   merges and deploys in the window; otherwise state the onset timestamp precisely
   so a human can check in minutes. A cause that predates the first failure is not
   the cause.
5. **Quantify the gap.** How many events were affected, over what period, for how
   many customers, and are any of them unrecoverable (a provider whose retention
   window has passed, an event with no idempotency key that cannot be safely
   replayed)? This number decides the severity and the urgency of the reply the
   support team owes people.
6. **Check whether it is still happening.** Distinguish an ongoing failure from a
   resolved one with an unprocessed backlog — they need different first actions.
   Say which it is in the first line of the report.
7. **Say what would fix it, and what would only appear to.** Separate the immediate
   recovery (replay this range, drain this dead-letter queue, restart this
   consumer) from the durable fix (persist before acknowledging, add idempotency
   keys, alert on dead-letter depth and consumer lag, monitor certificate expiry).
   Name every replay as a proposal with its risk — replaying non-idempotent events
   double-charges people — and never present one as done.

## Finding keys [#finding-keys]

The `key` names the *defect in the pipeline*, not this incident, so the same weak
hop keys identically the next time it fails and the ledger shows a genuinely
recurring fragility rather than a string of one-offs. Key on the component and the
failure mode.

* `webhook:stripe-events:signature-verification-failing`
* `queue:orders-sync:dead-letter-growth`
* `consumer:billing-worker:crash-loop`
* `endpoint:/hooks/provider:2xx-before-persist`
* `integration:zendesk-sync:silent-retry-exhaustion`
* `monitoring:dead-letter-depth:no-alert`

## Severity rubric [#severity-rubric]

Severity follows what is being lost right now, not how hard the fix is.

* **critical** — events are being lost unrecoverably, or the failure is ongoing and
  affects money, orders, or access for many customers.
* **high** — the failure is contained or stopped but a real backlog needs replaying,
  or events are being dropped silently (acknowledged and never persisted) so the
  loss is invisible without this investigation.
* **medium** — a recoverable backlog with retries still working, a fragile hop with
  no alerting, a single-customer failure with a known workaround.
* **low** — noisy but harmless retry patterns, cosmetic error handling, a
  dead-letter queue with old already-handled messages.
* **info** — the traced chain itself, recorded so the next incident of this shape
  starts from a map instead of from scratch.

## Output guidance [#output-guidance]

Open with one line answering the question the operator actually has:
is it still happening, what is being lost, and since when. Then a one-paragraph
executive summary naming the focus you investigated (and, when no focus was given,
which incident you chose and why). Then: a Delivery chain section walking sender →
edge → receiver → queue → consumer → application state, stating for each hop
whether the event was present, absent, or unverifiable and naming the system each
answer came from; a Change window section with the exact onset timestamp and the
candidate change; and an Impact section with the affected event count, customer
count, and what is unrecoverable. Close with two clearly separated lists —
Immediate recovery (each item marked as a proposal, with its risk, especially
replay of non-idempotent events) and Durable fix — followed by a recommendations
table (Action | Type | Risk | Owner). Never state or imply you replayed, re-drove,
or purged anything. Where a hop could not be checked because nothing is connected,
say which hop and what would be needed to see it.

<!-- generated by apps/server/scripts/export-checkups.ts, do not edit -->
