High Level Design
Idempotency in Distributed Systems
Why idempotency matters, how to implement idempotency keys in APIs and message queues, and the trade-offs of tracking duplicate requests.
Idempotency means performing the same operation multiple times results in the same final state as performing it once — even under retries, failures, or duplicates.
Why Idempotency Matters#
Distributed systems cannot guarantee exactly-once execution:
- At-least-once delivery is common (message queues, Kafka consumers)
- Failures occur between processing and acknowledgement
- Clients retry blindly on timeout
Without idempotency → correctness breaks.

Problems Solved#
- Duplicate requests due to: network retries, client timeouts, MQ redelivery
- Double side effects: double payments, duplicate orders, repeated emails/notifications
How It Works#

Flow#
- Client generates an Idempotency Key (e.g., UUID)
- Request sent to server with the key
- Server checks idempotency store (Redis, DB)
- If key exists → return the stored response (no reprocessing)
- If key is new → process the request
- Persist result + key atomically
Idempotency in APIs#
Typically applied to POST requests (which are non-idempotent by default).
Implementation:
- Idempotency-Key header in the request
- Key scoped to: user + operation type
- Same key with different payload → reject (conflict)
http
POST /payments
Idempotency-Key: a1b2c3d4-e5f6-7890-abcd-ef1234567890
Content-Type: application/json
{
"amount": 5000,
"currency": "USD",
"to": "user_456"
}
If the client retries with the same key, the server returns the original response without charging again.
Idempotency in Message Queues#
Messages may be delivered multiple times. The consumer must:
- Deduplicate via message ID — check a deduplication table or cache
- Or design logic to be inherently idempotent — e.g., upserts instead of inserts
Trade-offs#
| Pros | Cons |
|---|---|
| Strong correctness guarantees | Extra storage and lookups per request |
| Safe retries | Cleanup complexity (TTL on idempotency store) |
| Failure-tolerant systems | Needs careful key design |
| Slight latency overhead |