# DynamoDB capacity and table provisioning

> Over- and under-provisioned tables, on-demand versus provisioned fit, unused GSIs, and tables missing TTL.

Source: https://triagic.com/checkups/dynamodb-provisioning

## Scope [#scope]

A read-only capacity and cost review of the DynamoDB tables reachable through the
connected AWS credentials. Describe-and-measure only: no `UpdateTable`, no
`UpdateTimeToLive`, no deletes. Throttling is a correctness problem before it is
a cost problem, so it outranks savings everywhere in this procedure.

## Procedure [#procedure]

1. **Inventory.** `ListTables`, then `DescribeTable` on each. Record for every
   table: `BillingModeSummary` (PROVISIONED or PAY\_PER\_REQUEST),
   `ProvisionedThroughput` read/write units, `TableSizeBytes`, `ItemCount`,
   every global secondary index with its own key schema, projection type and
   throughput, and `StreamSpecification`. Also call `DescribeTimeToLive` per
   table. Note when the account has more tables than the credentials can describe
   and say so rather than reporting a partial inventory as complete.
2. **Actual consumption.** For each table and each GSI, pull CloudWatch metrics
   over at least 14 days (30 is better) at a 5-minute period from the
   `AWS/DynamoDB` namespace: `ConsumedReadCapacityUnits`,
   `ConsumedWriteCapacityUnits`, `ProvisionedReadCapacityUnits`,
   `ProvisionedWriteCapacityUnits`, `ReadThrottleEvents`,
   `WriteThrottleEvents`, `ThrottledRequests`,
   `SuccessfulRequestLatency`. Compute per table: mean consumed, p99 consumed,
   and mean-over-provisioned utilization for reads and writes separately — writes
   and reads are provisioned and billed independently and a table is regularly
   wrong on one and right on the other. If CloudWatch is not reachable through the
   connected servers, say so plainly: without consumption data this checkup can
   only report structure, not fit.
3. **Throttling first.** Any table or index with non-zero `ReadThrottleEvents` or
   `WriteThrottleEvents` in the window is under-provisioned *or* hot-keyed. Tell
   the two apart: throttling while table-level utilization sits well below 100%
   means the traffic is concentrated on one partition key, which more capacity
   will not fix — the fix is a write-sharding or key-design change. Throttling
   with utilization pinned near 100% is plain under-provisioning. Check whether
   auto-scaling is attached — that configuration is not in `DescribeTable`; it
   lives only in Application Auto Scaling
   (`DescribeScalableTargets` and `DescribeScalingPolicies` for the `dynamodb`
   service namespace). A provisioned table with no auto-scaling policy and any
   throttling at all is a standing incident risk, not a tuning note.
4. **On-demand versus provisioned fit.** Break-even utilization is set by the
   price ratio between an on-demand request unit and a provisioned capacity unit
   held for an hour. Derive that ratio from the currently published DynamoDB
   prices for the table's own region rather than a remembered number — AWS halved
   on-demand throughput pricing in November 2024 and any figure quoted in a
   procedure rots. As of that change the ratio is roughly 3–4x, putting break-even
   near 30% sustained utilization: below it on-demand is cheaper, above it
   provisioned is, and the band on either side of break-even is close enough that
   the change is not worth proposing. For each table compute the ratio of mean
   consumed capacity to peak consumed capacity:
   * PROVISIONED with mean utilization under \~30% of provisioned, or spiky traffic
     with long idle stretches → propose PAY\_PER\_REQUEST.
   * PAY\_PER\_REQUEST with steady, predictable consumption well above the
     break-even → propose PROVISIONED with auto-scaling, and give the target
     utilization (70% is the usual starting point).
     State the price ratio as approximate and region-dependent, say which prices and
     which date the break-even came from, and give the estimate as a range, never a
     false-precision dollar figure.
5. **Indexes.** For every GSI, compare its own `ConsumedReadCapacityUnits` and
   `ConsumedWriteCapacityUnits` against the base table's. A GSI with essentially
   zero consumed *read* capacity over the whole window but non-zero consumed write
   capacity is pure cost: it is being maintained on every base-table write and read
   by nobody. That is the single highest-confidence finding in this checkup. Also
   flag GSIs with `ProjectionType: ALL` on a wide item where only a few
   attributes are ever needed, and GSIs whose provisioned throughput is
   independently over- or under-set relative to their own consumption.
6. **Storage and lifecycle.** Compare `TableSizeBytes` growth across runs.
   `TableSizeBytes` and `ItemCount` are `DescribeTable` fields, not CloudWatch
   metrics, and DynamoDB refreshes them only about every six hours — so they are a
   run-over-run trend, never a live figure, and the report should say so. Any
   table holding event, session, log, audit or cache-shaped data with TTL disabled
   is growing forever and paying storage forever — check `DescribeTimeToLive` and
   flag it. Call `DescribeContinuousBackups` per table (PITR status is there, not
   in `DescribeTable`) and note tables with point-in-time recovery disabled where the
   data looks business-critical: that is a resilience finding, and it belongs in
   this report even though it costs money rather than saves it.
7. **Say what you could not see.** Region coverage, missing CloudWatch permissions,
   and tables the credentials could not describe all belong in the report. A
   capacity review that silently skipped half the account is worse than no review.

## Finding keys [#finding-keys]

The `key` must identify the *underlying issue* so the same problem lands on the
same ledger row next month rather than opening a fresh one — and so a fixed table
that regresses is detected as a regression rather than a new finding. Use
`<object-type>:<table-or-index-name>:<issue-slug>`. Use the real table name; for a
GSI, use `index:<table>.<index-name>:<issue>`. Never include capacity numbers,
dollar figures, or dates in the key.

* `table:orders:overprovisioned`
* `table:sessions:no-ttl`
* `table:events:throttled`
* `index:orders.gsi_status_created:unused`
* `table:audit_log:billing-mode-mismatch`

## Severity rubric [#severity-rubric]

* **critical** — sustained throttling on a table serving user-facing traffic, or a
  hot-partition pattern that more capacity cannot fix; requests are failing now.
* **high** — a billing-model or provisioning change worth more than 30% of this
  table's cost; a GSI consuming write capacity with no reads at all; a
  production-critical provisioned table with no auto-scaling and prior throttling.
* **medium** — steady over-provisioning in the 20–50% utilization band; an
  event/session/log-shaped table with no TTL; a wide `ALL` projection where a
  `KEYS_ONLY` or `INCLUDE` projection would serve every observed access.
* **low** — small tables provisioned above need where the absolute cost is trivial;
  cosmetic naming or tagging gaps; PITR disabled on non-critical data.
* **info** — capacity is well matched; growth is linear and explained; a new table
  appeared and is worth watching next run.

## Output guidance [#output-guidance]

Open with a one-paragraph executive summary: how many tables were
reviewed, how many are throttling, and the largest single capacity or billing-model
change available. Then sections for Throttling and risk, Capacity fit, Indexes, and
Storage and lifecycle — throttling first, because it is a correctness problem.
Close with a recommendations table (Table or index | Change | Expected effect |
Risk of the change | Effort), throttling fixes above savings. Give savings as
ranges and name the region-dependent pricing assumption behind them. Explicitly
list any table or region the credentials could not reach.

<!-- generated by apps/server/scripts/export-checkups.ts, do not edit -->
