Webhook and integration failure deep-dive
Run this against a named ticket or incident: trace a webhook or integration failure end to end and find where events are being lost.
Scope
A deep-dive triage procedure for events that were sent and never took effect: webhook deliveries, queue consumers, sync jobs, integration callbacks. Run on demand with a focus naming what to investigate. Read-only throughout — never replay, re-drive, re-queue, acknowledge, or purge anything, however tempting; say what should be replayed and let a human do it.
Procedure
- Start from the focus. The operator names a ticket or an incident — "ticket
#4821", "the Stripe webhook backlog from Tuesday", "orders stuck since the
deploy". Begin by loading it: pull that ticket, its body, its metadata and its
investigation history, and restate in one line what is claimed to be broken,
for whom, and since when. Everything after this is anchored to those three
facts. Use
tickets__search_similaron the focus text immediately — this exact failure has often happened before, and a resolved neighbour usually contains the answer and shortens the whole run. If no focus was given, do not stop. Fall back to finding the subject yourself: look for the most recent cluster of integration- or webhook-shaped tickets, or the most recent failed sync visible in the data available to you. State clearly at the top of the report which incident you selected and why, and invite the operator to re-run with an explicit focus if you picked wrong. - Establish the boundary conditions before tracing. Pin down the exact time window (first failure to now, or to the last known good event), the direction (inbound webhook the org receives, or outbound call the org makes), the volume (one event or thousands), and the blast radius (one customer or everyone). A single-customer failure and a total outage look identical in a ticket and need completely different investigations — get this wrong and the rest is wasted.
- Walk the delivery chain in order and find the last hop that has the event.
The chain is: sender → transport/edge → receiver endpoint → queue → consumer →
application state. Confirm the event's presence or absence at each hop before
moving to the next, and stop where it disappears — that hop is the fault, and
everything downstream is a symptom. Concretely, with whatever is connected:
- Sender — the provider's own delivery log or event list, with the response code it recorded. A provider showing 200s has done its job; a provider showing 4xx/5xx or retry exhaustion has just told you the answer.
- Transport and edge — CDN, proxy or gateway logs for the endpoint path in the window: TLS failures, 403s from a rule change, body-size or timeout limits. Signature-verification rejections live here or at the receiver and are a very common cause after a secret rotation.
- Receiver — application logs and error tracking around each failure timestamp. Look for the handler's own errors, and for a response that returned 2xx before durably persisting the event, which loses events silently and is the nastiest shape in this whole class.
- Queue — depth, age of the oldest message, consumer lag, dead-letter depth and in-flight counts. A growing dead-letter queue is the loudest possible evidence and is routinely never looked at.
- Consumer — is it running, is it erroring, is it crash-looping, was it scaled to zero, is it stuck on one poison message blocking a partition?
- Application state — query for the records the events should have created or updated and confirm the gap the customer is reporting is real. Do this last but always do it: it converts a plausible story into a verified one.
- Anchor to a change. Almost every failure of this class starts at a change. Ask what changed just before the first failure: a deploy, a config or secret rotation, a certificate expiry, a provider API version bump, an IP allowlist or firewall edit, a scaling change. Where source control or CI is connected, check merges and deploys in the window; otherwise state the onset timestamp precisely so a human can check in minutes. A cause that predates the first failure is not the cause.
- Quantify the gap. How many events were affected, over what period, for how many customers, and are any of them unrecoverable (a provider whose retention window has passed, an event with no idempotency key that cannot be safely replayed)? This number decides the severity and the urgency of the reply the support team owes people.
- Check whether it is still happening. Distinguish an ongoing failure from a resolved one with an unprocessed backlog — they need different first actions. Say which it is in the first line of the report.
- Say what would fix it, and what would only appear to. Separate the immediate recovery (replay this range, drain this dead-letter queue, restart this consumer) from the durable fix (persist before acknowledging, add idempotency keys, alert on dead-letter depth and consumer lag, monitor certificate expiry). Name every replay as a proposal with its risk — replaying non-idempotent events double-charges people — and never present one as done.
Finding keys
The key names the defect in the pipeline, not this incident, so the same weak
hop keys identically the next time it fails and the ledger shows a genuinely
recurring fragility rather than a string of one-offs. Key on the component and the
failure mode.
webhook:stripe-events:signature-verification-failingqueue:orders-sync:dead-letter-growthconsumer:billing-worker:crash-loopendpoint:/hooks/provider:2xx-before-persistintegration:zendesk-sync:silent-retry-exhaustionmonitoring:dead-letter-depth:no-alert
Severity rubric
Severity follows what is being lost right now, not how hard the fix is.
- critical — events are being lost unrecoverably, or the failure is ongoing and affects money, orders, or access for many customers.
- high — the failure is contained or stopped but a real backlog needs replaying, or events are being dropped silently (acknowledged and never persisted) so the loss is invisible without this investigation.
- medium — a recoverable backlog with retries still working, a fragile hop with no alerting, a single-customer failure with a known workaround.
- low — noisy but harmless retry patterns, cosmetic error handling, a dead-letter queue with old already-handled messages.
- info — the traced chain itself, recorded so the next incident of this shape starts from a map instead of from scratch.
Output guidance
Open with one line answering the question the operator actually has: is it still happening, what is being lost, and since when. Then a one-paragraph executive summary naming the focus you investigated (and, when no focus was given, which incident you chose and why). Then: a Delivery chain section walking sender → edge → receiver → queue → consumer → application state, stating for each hop whether the event was present, absent, or unverifiable and naming the system each answer came from; a Change window section with the exact onset timestamp and the candidate change; and an Impact section with the affected event count, customer count, and what is unrecoverable. Close with two clearly separated lists — Immediate recovery (each item marked as a proposal, with its risk, especially replay of non-idempotent events) and Durable fix — followed by a recommendations table (Action | Type | Risk | Owner). Never state or imply you replayed, re-drove, or purged anything. Where a hop could not be checked because nothing is connected, say which hop and what would be needed to see it.
Run it against your systems
This checkup is in the desktop app under Checkups. No card, read-only credentials you configure.