Skip to content

DynamoDB capacity and table provisioning

Over- and under-provisioned tables, on-demand versus provisioned fit, unused GSIs, and tables missing TTL.

Scope

A read-only capacity and cost review of the DynamoDB tables reachable through the connected AWS credentials. Describe-and-measure only: no UpdateTable, no UpdateTimeToLive, no deletes. Throttling is a correctness problem before it is a cost problem, so it outranks savings everywhere in this procedure.

Procedure

  1. Inventory. ListTables, then DescribeTable on each. Record for every table: BillingModeSummary (PROVISIONED or PAY_PER_REQUEST), ProvisionedThroughput read/write units, TableSizeBytes, ItemCount, every global secondary index with its own key schema, projection type and throughput, and StreamSpecification. Also call DescribeTimeToLive per table. Note when the account has more tables than the credentials can describe and say so rather than reporting a partial inventory as complete.
  2. Actual consumption. For each table and each GSI, pull CloudWatch metrics over at least 14 days (30 is better) at a 5-minute period from the AWS/DynamoDB namespace: ConsumedReadCapacityUnits, ConsumedWriteCapacityUnits, ProvisionedReadCapacityUnits, ProvisionedWriteCapacityUnits, ReadThrottleEvents, WriteThrottleEvents, ThrottledRequests, SuccessfulRequestLatency. Compute per table: mean consumed, p99 consumed, and mean-over-provisioned utilization for reads and writes separately — writes and reads are provisioned and billed independently and a table is regularly wrong on one and right on the other. If CloudWatch is not reachable through the connected servers, say so plainly: without consumption data this checkup can only report structure, not fit.
  3. Throttling first. Any table or index with non-zero ReadThrottleEvents or WriteThrottleEvents in the window is under-provisioned or hot-keyed. Tell the two apart: throttling while table-level utilization sits well below 100% means the traffic is concentrated on one partition key, which more capacity will not fix — the fix is a write-sharding or key-design change. Throttling with utilization pinned near 100% is plain under-provisioning. Check whether auto-scaling is attached — that configuration is not in DescribeTable; it lives only in Application Auto Scaling (DescribeScalableTargets and DescribeScalingPolicies for the dynamodb service namespace). A provisioned table with no auto-scaling policy and any throttling at all is a standing incident risk, not a tuning note.
  4. On-demand versus provisioned fit. Break-even utilization is set by the price ratio between an on-demand request unit and a provisioned capacity unit held for an hour. Derive that ratio from the currently published DynamoDB prices for the table's own region rather than a remembered number — AWS halved on-demand throughput pricing in November 2024 and any figure quoted in a procedure rots. As of that change the ratio is roughly 3–4x, putting break-even near 30% sustained utilization: below it on-demand is cheaper, above it provisioned is, and the band on either side of break-even is close enough that the change is not worth proposing. For each table compute the ratio of mean consumed capacity to peak consumed capacity:
    • PROVISIONED with mean utilization under ~30% of provisioned, or spiky traffic with long idle stretches → propose PAY_PER_REQUEST.
    • PAY_PER_REQUEST with steady, predictable consumption well above the break-even → propose PROVISIONED with auto-scaling, and give the target utilization (70% is the usual starting point). State the price ratio as approximate and region-dependent, say which prices and which date the break-even came from, and give the estimate as a range, never a false-precision dollar figure.
  5. Indexes. For every GSI, compare its own ConsumedReadCapacityUnits and ConsumedWriteCapacityUnits against the base table's. A GSI with essentially zero consumed read capacity over the whole window but non-zero consumed write capacity is pure cost: it is being maintained on every base-table write and read by nobody. That is the single highest-confidence finding in this checkup. Also flag GSIs with ProjectionType: ALL on a wide item where only a few attributes are ever needed, and GSIs whose provisioned throughput is independently over- or under-set relative to their own consumption.
  6. Storage and lifecycle. Compare TableSizeBytes growth across runs. TableSizeBytes and ItemCount are DescribeTable fields, not CloudWatch metrics, and DynamoDB refreshes them only about every six hours — so they are a run-over-run trend, never a live figure, and the report should say so. Any table holding event, session, log, audit or cache-shaped data with TTL disabled is growing forever and paying storage forever — check DescribeTimeToLive and flag it. Call DescribeContinuousBackups per table (PITR status is there, not in DescribeTable) and note tables with point-in-time recovery disabled where the data looks business-critical: that is a resilience finding, and it belongs in this report even though it costs money rather than saves it.
  7. Say what you could not see. Region coverage, missing CloudWatch permissions, and tables the credentials could not describe all belong in the report. A capacity review that silently skipped half the account is worse than no review.

Finding keys

The key must identify the underlying issue so the same problem lands on the same ledger row next month rather than opening a fresh one — and so a fixed table that regresses is detected as a regression rather than a new finding. Use <object-type>:<table-or-index-name>:<issue-slug>. Use the real table name; for a GSI, use index:<table>.<index-name>:<issue>. Never include capacity numbers, dollar figures, or dates in the key.

  • table:orders:overprovisioned
  • table:sessions:no-ttl
  • table:events:throttled
  • index:orders.gsi_status_created:unused
  • table:audit_log:billing-mode-mismatch

Severity rubric

  • critical — sustained throttling on a table serving user-facing traffic, or a hot-partition pattern that more capacity cannot fix; requests are failing now.
  • high — a billing-model or provisioning change worth more than 30% of this table's cost; a GSI consuming write capacity with no reads at all; a production-critical provisioned table with no auto-scaling and prior throttling.
  • medium — steady over-provisioning in the 20–50% utilization band; an event/session/log-shaped table with no TTL; a wide ALL projection where a KEYS_ONLY or INCLUDE projection would serve every observed access.
  • low — small tables provisioned above need where the absolute cost is trivial; cosmetic naming or tagging gaps; PITR disabled on non-critical data.
  • info — capacity is well matched; growth is linear and explained; a new table appeared and is worth watching next run.

Output guidance

Open with a one-paragraph executive summary: how many tables were reviewed, how many are throttling, and the largest single capacity or billing-model change available. Then sections for Throttling and risk, Capacity fit, Indexes, and Storage and lifecycle — throttling first, because it is a correctness problem. Close with a recommendations table (Table or index | Change | Expected effect | Risk of the change | Effort), throttling fixes above savings. Give savings as ranges and name the region-dependent pricing assumption behind them. Explicitly list any table or region the credentials could not reach.

Run it against your systems

This checkup is in the desktop app under Checkups. No card, read-only credentials you configure.