# Spending and budget

> Setting a monthly LLM cap, what happens when it is reached, and where to see the running total.

Source: https://triagic.com/docs/admin/spending

**Organization → Spending** sets one number: a **monthly LLM cap in USD**.

Set it before anyone starts running investigations. It takes ten seconds and it's the
only backstop against a prompt loop that costs money.

## Setting the cap [#setting-the-cap]

Type a positive number and save. Leave the field blank and save to remove the cap
(unlimited).

The current cap is shown above the field. A blank input with a cap already set is not
"unlimited" until you actually save it.

## What happens at the cap [#what-happens-at-the-cap]

The cap is enforced on the desktop, in three places, and it fails gracefully in each:

**Background triage** checks the cap before starting a run. A ticket for an over-budget
organization is **deferred and retried later with backoff**, not failed. Nothing is
lost; work resumes when the cap resets or is raised. Only that organization's tickets
are affected, never the whole queue.

**Console and ticket chat** check before opening the response stream, so an over-cap
request gets a plain refusal (with the cap and the month-to-date spend in the message)
rather than half a streamed answer.

**A long multi-turn run** re-checks on every API call, so a run that crosses the cap
mid-flight is stopped rather than allowed to finish.

## Per-run cap for generation [#per-run-cap-for-generation]

Generating a [knowledge doc](/docs/admin/knowledge#generate-a-doc-from-a-prompt) or
[playbook](/docs/admin/playbooks#generate-a-playbook-from-a-prompt), revising one,
and proposing, building, changing or repairing a [dashboard](/docs/desktop/dashboards)
each run as one AI run. Those five kinds of run have a second, smaller limit: a hard
cap on what any single run may spend. It does not apply to triage, the Console,
checkups or reports.

The default is **$1.00**. Admins change it under **Organization → Settings**, in
**Hard cap per generation run**, next to the monthly cap: $0.50, $1, $2, $5, or **No
cap**. Desktops pick up the change on their next config sync. An install that isn't
centrally managed uses the $1.00 default.

The desktop checks the cap before every model call in the run. A run that reaches it
stops there and keeps what it already finished: a dashboard build keeps the charts it
had saved, and the rest are marked as not built. A run that had finished nothing
fails with &#x2A;"Stopped at the $1.00 per-run cap."* (with your cap in place of $1.00).
Generation runs also count toward the monthly cap like any other AI use.

## Prices on buttons [#prices-on-buttons]

Every button that starts one of those runs shows a price first, such as **Build 6
items \~$0.40**, and the form beside it shows the estimate, the month's spend against
the monthly cap and the per-run cap. The estimate comes from your organization's
recent runs of the same kind: once there are at least three, it's the middle of
their costs. Before that, and always for dashboard builds, it's the model's price for
a typical run (times the number of charts, for a build).

Controls that never spend AI say **Free**: publishing a doc or playbook, editing a
draft or a saved query by hand, and everything a member does on a dashboard.

## Viewing dashboards is free [#viewing-dashboards-is-free]

A dashboard costs AI once, when an admin builds it, and again only when an admin
adds, changes or repairs a chart. Opening it, changing the date range and refreshing
it run its saved queries directly against your data sources with no model, so they
never count toward either cap and never appear on the Usage page.

## Where the spend total is [#where-the-spend-total-is]

Not on this card, which sets the cap and nothing else. Each member's install meters
against the cap locally, and also pushes a daily rollup to the cloud on its regular
sync tick, so the portal's **Usage** page shows the org-wide running total across every
machine: month-to-date spend, a per-model, per-task and per-member split, and a daily
cost chart. That page is admin- and owner-only.

What leaves a desktop is aggregate figures only: per day, member, task and model, the
number of calls, the prompt and completion token counts, and the cost in USD, plus the
machine's id. Prompts, responses, ticket content and individual call records never go
up. Each push re-sends the last 35 days whole.

Per-model p50/p95 latency and the 50 most recent calls with links back to what caused
them exist only on the desktop's **Usage** page. See
[History and usage](/docs/desktop/history-and-usage#usage).

## Picking a number [#picking-a-number]

There is no universally right cap, but there is a decent method:

1. Set a deliberately low cap for the first week, low enough that you expect to hit
   it.
2. Look at the desktop **Usage** page's per-task split at the end of it.
3. Multiply out for your real ticket volume and set the real cap with headroom.

The per-task split is the useful part. It is common for classification (which runs on
every incoming ticket, whether or not anyone reads the result) to be a larger share
than expected, and that is a signal to reduce the number of enabled playbooks rather
than to raise the cap.

## Levers other than the cap [#levers-other-than-the-cap]

The cap stops spending. These reduce it:

* **Narrow playbook data sources.** Fewer systems means fewer agent iterations.
* **Sharpen triage instructions.** A playbook that names the table to check first
  finishes in fewer turns than one that says "investigate".
* **Prune playbooks.** Every enabled playbook is a candidate the classifier evaluates.
* **Choose the model deliberately.** A cheaper model that needs three times the
  iterations is not cheaper.

## Audit [#audit]

Changing the cap writes an `org.update` audit entry with the before and after values.
