The situation
Corvane Logistics runs four warehouses for fourteen brands. Every brand sends its shipping orders as EDI 940 messages, and a message that fails to sync is an order the warehouse never sees. Nobody picks it, nobody ships it, and the brand finds out from its own customer.
The Warehouse floor dashboard already had a KPI for it: Failed syncs (last hour), a saved read-only query that counts rejected messages in edi_messages. On a normal night it reads between 2 and 5. Nobody looks at a dashboard at three in the morning.
So the ops lead pressed Alert on that tile and typed: page the warehouse ops team if failed syncs go over 10. The button showed its price before the ops lead pressed Write rule. One model call turned the sentence into one rule, and the rule editor opened to confirm it:
- Name: Failed syncs over 10.
- Condition: Failed syncs (last hour) > 10, on one check, because a failure count needs no second opinion.
- Severity: critical.
- Team: Warehouse ops.
Saving it was free. From then on the rule is a comparison, not a prompt.
What Triagic looked at
Every 15 minutes the rule owner's desktop ran the tile's saved query against MongoDB, read-only, and compared the number with 10. That is the whole check. No model is called, so a check costs nothing.
The rule lives on the desktop in the shift office, which stays on overnight. That matters: checks run on the rule owner's desktop, and only while the app is running. The Rules tab shows "last checked" on every rule, so a quiet rule and a stopped one never look the same.
What it found
The checks from 01:12 to 02:57 read 2, 3, 2, 4, 3, 5, 4 and 7. At 03:12 the check read 23, against a threshold of 10, and the rule fired critical. The ladder for critical is fixed, and it ran in this order:
- The owner's desktop: a bell entry and a notification that stays on screen until someone clicks it.
- Email, to the owner and every member of Warehouse ops: the rule, the dashboard,
Value: 23 (threshold 10), the time it opened, and an Acknowledge link, which Corvane's emails carry because the install is cloud-managed.
- Slack, in the alert channel an admin set on the Alerts page: the same rule and value, with the same link.
A critical alert keeps paging every 15 minutes, up to 8 times, until someone acknowledges it. Ana, on call for warehouse ops, opened the email on her phone at 03:19, signed in to the web portal and pressed Acknowledge. At 03:27 the desktop checked with the cloud before sending the first repeat, found her acknowledgement there, and sent nothing. That check read 31; the 03:42 one read 29. The event stayed acknowledged and open.
The record on the Alerts page is the evidence: the severity, 23 vs 10, who acknowledged it and when, and a note for any channel that failed. The dashboard name links to the tile, and Query on the tile shows the saved query that produced the 23. Nothing in that chain is a model's opinion.
What changed
At 07:30 the ops lead asked the Console the morning question: which brand are the failed syncs from, and which orders are inside them? It read edi_messages in MongoDB and checked orders in Postgres, read-only, with both queries on screen. Every rejection since 02:46 was a Saltgrass Apparel 940, rejected with DUPLICATE_CONTROL_NUMBER: their system had reused control numbers Corvane had already processed, so new shipping orders bounced as repeats. 54 messages, 135 orders, none of them in the warehouse system.
Saltgrass resent the batch with fresh control numbers at 08:15, and the orders made the 09:00 wave in Reno. The rule resolved on its own once the hour's count dropped back under 10, and everyone who had been paged got a short "resolved" message on the same channels.
The night shift found out at 03:12 instead of at the 07:00 handover. Build the dashboard first (Monday dashboard), then read the alerts docs for the full severity ladder.