Query vs Scan in DynamoDB
Query reads a single by partition key (optionally narrowing the
sort key); Scan reads the whole table and filters afterwards. They look
similar in the API but they bill — and scale — completely differently.
When should I use Query vs Scan in DynamoDB?
Use Query whenever you can name the partition you need — it reads one and bills only for matched items. Reach for Scan only for one-off exports or tiny tables; it reads every item and bills the whole table before any FilterExpression runs. On real data, Query wins.
- Query is targeted: you pay for the items in the matched partition.
- Scan is exhaustive: you pay to read every item, then throw most away with a
FilterExpressionthat runs after the read is metered.
On a table of any real size, a Scan with a filter is the classic "why is my
bill huge and my latency worse than RDS" footgun.
Side by side
| Query | Scan | |
|---|---|---|
| Reads | One partition (by PK) | Every item in the table |
| Billed capacity | Items matched in the partition | Whole table, before filtering |
FilterExpression | Applied after the read — still billed for the read | Same — filtering never cuts cost |
| Latency | Flat as the table grows | Grows with table size |
| Pagination | 1 MB/page → LastEvaluatedKey | 1 MB/page; parallelisable |
| Use it for | Known access patterns | One-off exports, tiny config tables |
The key trap: a FilterExpression runs after DynamoDB meters the read, on
both operations. A Scan that "returns 10 rows" can bill for reading a million —
filtering is a convenience, never a cost control.
What a full Scan actually costs
Put numbers on it. DynamoDB meters reads in 4 KB units: a
read costs one
per 4 KB, an
read half that. Query and
Scan sum the size of every item they touch — not every item they return —
and round up to the next 4 KB.
Take a 1-million-item table averaging 2 KB per item (~2 GB of data), and an access pattern that needs 10 of those items:
| Items read | Data metered | Read units (eventually consistent) | |
|---|---|---|---|
Scan + FilterExpression | 1,000,000 | ~2 GB | ~262,000 |
Query on a matching key | 10 | 20 KB | 3 |
Same 10 items, five orders of magnitude apart — and the Scan bills that
~262,000 every time it runs, whether the filter matches ten items or none. On
billing those are request units straight onto the
bill; on tables a big Scan competes with
production traffic for throughput and can throttle it into
ProvisionedThroughputExceededException.
Three more cost facts that surprise people:
Select: COUNTis not free. A counting Query or Scan consumes exactly the same read capacity as reading the items — it just doesn't return them.Limitcaps items evaluated, not items matched. Combined with a filter, a page can come back empty while still billing a full page of reads.- You never have to guess. Pass
ReturnConsumedCapacity: TOTALand every response reports the capacity it just consumed.
Check what your own items weigh with the item size calculator, then turn read units into a monthly bill with the pricing calculator.
Use Query
Query PK = "USER#42" AND SK begins_with "ORDER#"
If you find yourself reaching for Scan to answer a common access pattern, that
is a modelling signal: add a Global Secondary Index so the
pattern becomes a Query.
The choice comes down to one question — can you name the partition you need?
If the key is known you Query; if not, add a GSI to make it one, and fall back
to Scan only when no key fits.
When Scan is fine
One-off exports, tiny config tables, and background jobs that page through the
whole table deliberately. Use Segment/TotalSegments to split a Scan across
workers (a — see
parallel Scans in DynamoDB) when you genuinely
must read everything, and page it properly with LastEvaluatedKey
(pagination guide). If a Scan you already run is the
problem, why Scan is slow and expensive
walks the triage.
A reflexive SELECT * FROM table over DynamoDB is the same anti-pattern in
PartiQL clothing — it compiles to a Scan. When you really do need cross-item
analytics (a GROUP BY, a JOIN, an aggregate), DynoTable's SQL Workbench runs
them client-side over a bounded result set instead of hammering the table.
Try DynoTable to run and inspect these queries against your own tables — it shows the consumed capacity of every operation it runs.