# PII exposure sweep

> Personal data sitting in tickets, chat threads and reachable data stores: where it is, how much, and what should be redacted.

Source: https://triagic.com/checkups/pii-exposure

## Scope [#scope]

Find personal data that has accumulated where it should not be: pasted into
support tickets, quoted into chat threads and reports, or sitting in a store any
agent run can reach. This is a locate-and-count exercise; it changes nothing and
redacts nothing.

**The report must not become the incident.** Never quote a raw personal
identifier, never reproduce a card number, national id, password, token or full
email address, and never copy a matched string into `detail` or `evidence`.
Describe instead: the kind of data, where it sits (ticket id, field, table and
column), how many instances, and a masked shape (`4111-****-****-1111`,
`j***@example.com`). A finding that requires the reader to see the value to act
on it is a finding written wrongly.

## Procedure [#procedure]

1. **Decide what counts.** Work with an explicit list and state it in the report:
   payment card numbers, bank account and routing numbers, national identifiers
   (SSN, NI, PAN, Aadhaar and equivalents), passport and licence numbers, full
   postal addresses, dates of birth, phone numbers, email addresses, precise
   geolocation, health information, and authentication material (passwords, API
   keys, tokens, session cookies, private keys). Treat authentication material as
   the highest tier regardless of volume — one leaked key is worse than a thousand
   email addresses.
2. **Sweep the support corpus.** Using the ticket data available to you, look
   across ticket subjects, bodies, and their investigation write-ups for the
   categories above. `tickets__search_similar` is a good way to find the
   *clusters* — search for the shapes that recur ("card ending", "my password is",
   "attached is my ID", "here is my token") and follow them to the tickets. Pay
   particular attention to: customers pasting card or identity documents into a
   description; support agents pasting production rows into a reply; credentials
   pasted "just to reproduce it". Count instances per category and note which
   tickets are still open (open tickets keep spreading the data through
   notifications and replies).
3. **Sweep the agent's own output.** Investigations, reports and chat threads
   quote tool results, and a query result pasted into a thread persists far longer
   than the query did. Look for personal data that entered through Triagic's own
   working, not through the customer — that is a finding against the team, and it
   is the one most likely to be fixable by changing a playbook's instructions.
4. **Assess reachable stores, where any are connected.** Do not dump data. Read
   *schemas*: table and column names across the connected databases and search
   indices, and flag columns whose names denote personal data (`ssn`, `dob`,
   `card_number`, `passport`, `address_line1`, `phone`, `email`,
   `full_name`, `ip_address`). For each, note whether the column looks
   encrypted or tokenized from its type and sample-free metadata, and whether the
   connected role can read it at all. The question this step answers is: &#x2A;what
   would an agent run be able to retrieve if a ticket asked it to?* Where sampling
   a value is unavoidable to confirm a column is not already tokenized, report only
   the shape, never the value. If no store is connected, say the assessment covered
   the support corpus only.
5. **Retention.** Personal data that had a reason to exist two years ago rarely
   still does. Note the age distribution of what you found: PII in tickets closed
   more than a year ago is a retention finding regardless of how it got there, and
   it is usually the cheapest to remediate in bulk.
6. **Spread.** For each cluster, say where else it has travelled: quoted in a
   reply, copied into a linked ticket, included in a scheduled report, exported.
   Volume matters less than reach.

## Finding keys [#finding-keys]

The `key` identifies the *location and kind of exposure* — never the data itself,
because a key is stored and displayed. Use
`<surface>:<location>:<data-kind>`, where the location is a stable address
(a table and column, a ticket id, a playbook name) and never a value or a hash of
one. The same cluster must key identically next week so the ledger shows whether a
redaction actually happened.

* `tickets:body:payment-card`
* `ticket:4821:credentials-pasted`
* `store:postgres.public.customers.ssn:unmasked-column`
* `thread:playbook-billing-report:query-result-with-emails`
* `retention:closed-tickets-over-1y:personal-data`

## Severity rubric [#severity-rubric]

* **critical** — authentication material (password, API key, private key, session
  token) sitting in any ticket, thread or report; unmasked payment card numbers or
  national identifiers in a corpus that is broadly readable.
* **high** — a recurring pattern rather than an incident: a workflow that routinely
  puts identity documents or card data into tickets; an unmasked personal-data
  column readable by an integration role that every agent run can reach.
* **medium** — contact-level personal data (emails, phone numbers, addresses)
  accumulating beyond what support actually needs; personal data quoted into
  Triagic's own reports by a playbook's instructions.
* **low** — small, aged, low-sensitivity instances in closed tickets; a
  personal-data column that is already tokenized but poorly named.
* **info** — the sweep's coverage and counts, recorded as the baseline the next run
  compares against.

## Output guidance [#output-guidance]

Open with a one-paragraph executive summary: which categories were
searched, the highest-severity category actually found, and the total count by
category — counts and categories only, never examples. Then sections for Support
corpus, Triagic's own output, Reachable stores, and Retention. Every location must
be precise enough to act on (ticket id, table and column, playbook name) and every
value must be masked or described rather than reproduced. Close with a
recommendations table (Location | Data kind | Recommended action | Who acts |
Urgency), redaction and rotation first, then the process change that stops the
pattern recurring — a one-off cleanup with no process change will simply re-emit
the same finding next week. State plainly which surfaces were not searched.

<!-- generated by apps/server/scripts/export-checkups.ts, do not edit -->
