Idempotency Keys: Safe API Retries in FastAPI

A network retry can submit the same request twice. How idempotency keys in FastAPI make POST endpoints safe to retry without duplicate side effects.

A client sends POST /requests, and the handler runs. Somewhere between the database commit and the response reaching the browser, the connection drops. The client saw a timeout, not a 201, so it does the reasonable thing and retries. Now the same action has run twice. On an ops console, that means two copies of a production request queued into the grid. On a payments API, it means charging a card twice.

I hit this while building the operator console for CMS workflow operations at CERN. Operators trigger actions that clone requests, change site policies, and re-queue work across the Worldwide LHC Computing Grid. A double-submit there is not cosmetic. It schedules real compute.

The fix is old and well understood: idempotency keys. This post is for backend engineers who have a POST endpoint with side effects and need a retry to be safe, whether the retry comes from a nervous user double-clicking, a load balancer, or a client library with automatic backoff. I cover how the keys work, why a simple cache is not enough, an atomic implementation in Postgres and FastAPI, and where it breaks in production.

What “idempotent” actually buys you

An operation is idempotent if doing it twice leaves the system in the same state as doing it once. The HTTP spec defines GET, PUT, and DELETE as idempotent. POST is not, and that is the whole problem: POST is where you create things, and creating the same thing twice is exactly the failure you want to avoid.

You cannot make POST idempotent by wishing. What you can do is attach a client-generated token to each logical operation, so the server can recognise a repeat and refuse to do the work a second time. That token is the idempotency key. The client sends it in a header, by convention Idempotency-Key (now written up in an IETF draft), and reuses the same value on every retry of that one operation.

Stripe’s API popularised this pattern. Their docs are worth reading even if you never touch payments, because they document the edge cases most homegrown versions miss.

The key belongs to the operation, not the request. Two different button presses are two operations and get two keys. Five retries of one button press are one operation and share one key. Getting that boundary right is the client’s job; the server just trusts the key.

How the mechanism works

The server keeps a small record for each key it has seen. The first request with a given key does the work and saves its response against the key. Any later request carrying the same key gets that saved response replayed, and the side effect never runs again.

A sequence diagram with three lanes: client, API handler, and key store. The first POST carries Idempotency-Key k-91; the handler inserts k-91 into the store as a miss and claims it, runs the side effect exactly once, stores the response body and status, and returns 201 with id 4471. A network hiccup means the client never saw the 201, so it retries the identical request. The handler selects k-91, gets a hit, runs no side effect, and replays the stored 201 with id 4471.

That is the happy path, and it really is simple. The interesting engineering is in the record itself, because a naive version has a race condition that shows up the first time two retries arrive close together.

Why a plain cache is not enough

The tempting implementation goes like this: look up the key; if it is there, return the stored response; otherwise, run the handler and store the result. That is a check-then-act sequence, and check-then-act across two requests is a race.

Picture a client that fires a request, does not see a response fast enough, and fires the retry while the first is still running. Both requests look up the key, both miss, and both run the handler. You now have the duplicate you were trying to prevent, and you built a whole caching layer to still get it. On a busy service this is not a rare timing window, because aggressive client retry policies are designed to fire quickly.

So the record needs three states, not two, and the move into the “claimed” state has to be atomic: one indivisible step that no other request can slip into.

A state diagram for one idempotency record. It starts as no record, meaning the key was never seen. The first request atomically inserts the row and moves it to in_progress, where the row is claimed and the work is running. When the work finishes, the response body and status are saved and the state becomes completed. A retry that arrives while the record is in_progress gets a 409 Conflict telling it to try again shortly. A retry that arrives after completion gets the stored response replayed with no work. After a TTL the row expires and the key is free again. A note says a payload arriving with a used key but a different body should be rejected with 422, never silently served the old answer.

The in_progress state is the part people skip and then debug in production. It means “someone is already doing this work.” It exists so a concurrent retry has something to collide with, instead of quietly starting a second copy.

Making the claim atomic in Postgres

The cleanest way to claim a key atomically is to let the database enforce it. Put a unique constraint on the key, and use an insert that fails loudly on collision. In Postgres, INSERT ... ON CONFLICT does exactly this in a single statement, so there is no window between checking and claiming. The table below makes the key its primary key and stores the request hash, the state, and the saved response:

CREATE TABLE idempotency_keys (
    key            TEXT PRIMARY KEY,
    request_hash   TEXT NOT NULL,
    status         TEXT NOT NULL DEFAULT 'in_progress',
    response_code  INT,
    response_body  JSONB,
    created_at     TIMESTAMPTZ NOT NULL DEFAULT now()
);

The request_hash column matters more than it looks. A key identifies an operation, but a buggy or malicious client could reuse a key with a different body. If you blindly replayed the stored response, you would answer the wrong question. Storing a hash of the request lets you detect that mismatch and reject it instead.

Wiring it into FastAPI as a dependency

FastAPI’s dependency system is the right seam for this. The idempotency logic is cross-cutting, and copy-pasting it into every mutating handler is how the copies drift out of sync. A dependency lets you attach it declaratively to the routes that need it.

Here is the claim step, which runs before the handler body. It tries to insert the key; if the key already exists, it reads the existing record and decides how to answer:

import hashlib, json
from fastapi import Request, HTTPException, Header

async def claim_key(
    request: Request,
    idempotency_key: str = Header(..., alias="Idempotency-Key"),
):
    body = await request.body()
    request_hash = hashlib.sha256(body).hexdigest()

    async with request.app.state.db.acquire() as conn:
        row = await conn.fetchrow(
            """
            INSERT INTO idempotency_keys (key, request_hash)
            VALUES ($1, $2)
            ON CONFLICT (key) DO NOTHING
            RETURNING key
            """,
            idempotency_key, request_hash,
        )

        if row is not None:
            # We won the insert. This request owns the work.
            return {"key": idempotency_key, "fresh": True}

        # Key already exists, so load the existing record.
        existing = await conn.fetchrow(
            "SELECT request_hash, status, response_code, response_body "
            "FROM idempotency_keys WHERE key = $1",
            idempotency_key,
        )

    if existing["request_hash"] != request_hash:
        raise HTTPException(422, "Idempotency-Key reused with a different body")

    if existing["status"] == "in_progress":
        # A twin request is still running. Tell the client to back off.
        raise HTTPException(409, "Request with this key is still processing")

    # Completed: replay the stored response.
    raise ReplayResponse(existing["response_code"], existing["response_body"])

ON CONFLICT (key) DO NOTHING ... RETURNING key is the whole trick:

  • If the insert succeeds, RETURNING hands back the row, and this request is the owner.
  • If the key already existed, DO NOTHING suppresses the row and fetchrow returns None, so the code falls through to the read path. There, a different body gets a 422, a record still in_progress gets a 409, and a completed record is replayed.

Only one concurrent request can ever win the insert, because the primary key constraint serialises them at the database. There is no application-level lock and no Redis SETNX dance. The constraint you already have does the work.

The handler then does its job and writes the response back against the key before returning:

@app.post("/requests", dependencies=[Depends(claim_key)])
async def create_request(payload: RequestIn, claim=Depends(claim_key)):
    result = await schedule_workflow(payload)     # the real side effect

    async with app.state.db.acquire() as conn:
        await conn.execute(
            "UPDATE idempotency_keys "
            "SET status='completed', response_code=201, response_body=$2 "
            "WHERE key=$1",
            claim["key"], json.dumps(result),
        )
    return result

ReplayResponse is a small custom exception, with a handler registered via @app.exception_handler that turns the stored code and body back into a real response. Using the exception path keeps the replay out of the normal handler, so on a repeat the side-effect code is never even reached.

Where it bites

The skeleton above is maybe forty lines. The reasons it goes wrong in production are more interesting than the code, and most of them are about what happens when the first request does not finish cleanly.

A crash between the side effect and the response write leaves a stranded in_progress row. If schedule_workflow succeeds but the process dies before the UPDATE, the key is stuck in in_progress forever, and every retry gets a 409. This is the hardest case, because it forces a question you cannot dodge: is your side effect itself idempotent? If scheduling the same workflow twice is safe, you can let a stuck key expire and be retried. If it is not, you need the side effect and the key update in the same database transaction, so they commit or roll back together. That only works when the side effect is a database write. An external API call cannot join your transaction, which is the real reason distributed idempotency is hard.

Keys need a TTL, and the TTL is a real tradeoff. You cannot keep every key forever. Expire them too fast, and a legitimate slow retry (a mobile client on a bad connection coming back after two minutes) misses the record and re-runs the work. Expire them too slowly, and the table grows without bound. Stripe keeps keys for 24 hours, which is a sane default for user-facing retries. Sweep expired rows on a schedule. This is the same lifecycle-of-a-record thinking behind OpenSearch index expiry, just at a different scale.

The 409 is a signal, not an error. A client that gets a 409 should wait and retry with the same key, because the work is genuinely still running. That is exactly the retry-with-backoff behaviour I have written about for rate limits, pointed at a different status code. If your clients treat 409 as fatal, the pattern does not help them.

A replayed response can be stale. The record captures the response as it was the first time. If the underlying resource changed between the original request and the retry, the replayed body is out of date. For a create-style endpoint that returns an ID, this is fine, because the ID does not change. For anything returning mutable state, decide deliberately whether the client wants “what happened when you asked” or “what is true now.”

Connection pool pressure is easy to overlook. The claim adds an extra database round trip to every write request, and it holds a connection while the handler runs. On a busy endpoint, that changes your pool math. Read up on async connection pooling before you turn this on for a high-traffic route, because a too-small pool turns your safety feature into a bottleneck.

What I would keep and what I would change

The Postgres-native approach earns its place when the side effect is also a Postgres write. Then the claim and the effect share one transaction, and the stranded-in_progress problem disappears. That is the version I would reach for first on the CMS console, where most operator actions land in the same database anyway.

If the side effect is an external call (a payment, or a message to another service), I would stop pretending a single transaction solves it. Instead, I would reach for an outbox-style pattern: record the intent transactionally, then have a separate worker perform the external call with its own idempotency guarantees and mark it done.

Idempotency keys at the edge and idempotency at the effect are two different layers. Conflating them is how you end up with a system that is safe against retries in the demo and duplicates work in production. Pair this with a circuit breaker on the external dependency, and you have a request path that degrades predictably instead of amplifying every retry storm into duplicate work.

Idempotency is unglamorous, and that is the point. It is the difference between an ops action you can retry without thinking and one where every timeout makes an operator wonder whether they just scheduled the same grid job twice. On a system where the side effects cost real compute, that peace of mind is the feature.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.