Spending and budget
Setting a monthly LLM cap, what happens when it is reached, and where to see the running total.
Organization → Spending sets one number: a monthly LLM cap in USD.
Set it before anyone starts running investigations. It takes ten seconds and it's the only backstop against a prompt loop that costs money.
Setting the cap
Type a positive number and save. Leave the field blank and save to remove the cap (unlimited).
The current cap is shown above the field. A blank input with a cap already set is not "unlimited" until you actually save it.
What happens at the cap
The cap is enforced on the desktop, in three places, and it fails gracefully in each:
Background triage checks the cap before starting a run. A ticket for an over-budget organization is deferred and retried later with backoff, not failed. Nothing is lost; work resumes when the cap resets or is raised. Only that organization's tickets are affected, never the whole queue.
Console and ticket chat check before opening the response stream, so an over-cap request gets a plain refusal (with the cap and the month-to-date spend in the message) rather than half a streamed answer.
A long multi-turn run re-checks on every API call, so a run that crosses the cap mid-flight is stopped rather than allowed to finish.
Per-run cap for generation
Generating a knowledge doc or playbook, revising one, and proposing, building, changing or repairing a dashboard each run as one AI run. Those five kinds of run have a second, smaller limit: a hard cap on what any single run may spend. It does not apply to triage, the Console, checkups or reports.
The default is $1.00. Admins change it under Organization → Settings, in Hard cap per generation run, next to the monthly cap: $0.50, $1, $2, $5, or No cap. Desktops pick up the change on their next config sync. An install that isn't centrally managed uses the $1.00 default.
The desktop checks the cap before every model call in the run. A run that reaches it stops there and keeps what it already finished: a dashboard build keeps the charts it had saved, and the rest are marked as not built. A run that had finished nothing fails with "Stopped at the $1.00 per-run cap." (with your cap in place of $1.00). Generation runs also count toward the monthly cap like any other AI use.
Prices on buttons
Every button that starts one of those runs shows a price first, such as Build 6 items ~$0.40, and the form beside it shows the estimate, the month's spend against the monthly cap and the per-run cap. The estimate comes from your organization's recent runs of the same kind: once there are at least three, it's the middle of their costs. Before that, and always for dashboard builds, it's the model's price for a typical run (times the number of charts, for a build).
Controls that never spend AI say Free: publishing a doc or playbook, editing a draft or a saved query by hand, and everything a member does on a dashboard.
Viewing dashboards is free
A dashboard costs AI once, when an admin builds it, and again only when an admin adds, changes or repairs a chart. Opening it, changing the date range and refreshing it run its saved queries directly against your data sources with no model, so they never count toward either cap and never appear on the Usage page.
Where the spend total is
Not on this card, which sets the cap and nothing else. Each member's install meters against the cap locally, and also pushes a daily rollup to the cloud on its regular sync tick, so the portal's Usage page shows the org-wide running total across every machine: month-to-date spend, a per-model, per-task and per-member split, and a daily cost chart. That page is admin- and owner-only.
What leaves a desktop is aggregate figures only: per day, member, task and model, the number of calls, the prompt and completion token counts, and the cost in USD, plus the machine's id. Prompts, responses, ticket content and individual call records never go up. Each push re-sends the last 35 days whole.
Per-model p50/p95 latency and the 50 most recent calls with links back to what caused them exist only on the desktop's Usage page. See History and usage.
Picking a number
There is no universally right cap, but there is a decent method:
- Set a deliberately low cap for the first week, low enough that you expect to hit it.
- Look at the desktop Usage page's per-task split at the end of it.
- Multiply out for your real ticket volume and set the real cap with headroom.
The per-task split is the useful part. It is common for classification (which runs on every incoming ticket, whether or not anyone reads the result) to be a larger share than expected, and that is a signal to reduce the number of enabled playbooks rather than to raise the cap.
Levers other than the cap
The cap stops spending. These reduce it:
- Narrow playbook data sources. Fewer systems means fewer agent iterations.
- Sharpen triage instructions. A playbook that names the table to check first finishes in fewer turns than one that says "investigate".
- Prune playbooks. Every enabled playbook is a candidate the classifier evaluates.
- Choose the model deliberately. A cheaper model that needs three times the iterations is not cheaper.
Audit
Changing the cap writes an org.update audit entry with the before and after values.