The situation
11:40. An email from ops@arcadialabs.io: their bulk import has been stuck at 82% for two hours, it's the third time this quarter, and if it isn't fixed today they're "evaluating alternatives".
The usual first reply is an apology and a promise to look into it. The customer knows what that means: more waiting. And the agent writing it can't look, because the import worker, the jobs table and the logs all sit behind engineering access.
What Triagic looked at
The ticket was triaged when it arrived, before anyone opened it. The investigation ran read-only against the systems we'd connected:
- Kubernetes: pod status and recent events for the
import-worker deployment.
- Postgres: the customer's row in
import_jobs, inside a read-only transaction.
- OpenSearch: the worker's logs for the window around the stall.
Personal data in the ticket (the email address, phone numbers, IDs) is redacted before any of it reaches the model.
What it found
The Root cause column in the ticket queue already read: import worker OOM-killed on a 240 MB CSV. The ticket detail had the evidence behind it:
kubernetes: import-worker OOMKilled at 11:31.
postgres: import_jobs row 8412 stalled at 82%.
opensearch: a heap error while parsing a 240 MB file.
The suggested reply was drafted from those facts: what broke, the workaround (split the file) and that a fix for large files is going to engineering. The agent edited one line and sent it from the ticket.
What changed
First reply: 11 minutes after the angry email, with the cause in it. Escalations to engineering: none. The customer's answer was "ok, appreciate the detail. we'll split the file."
Engineering heard about it once, as a filed issue with the root cause and evidence attached, not as an interrupt in a chat channel. Read how tickets are triaged in the docs, or connect Kubernetes and Postgres and try it on your own queue.