how to tell if an engineer really understands idempotency

by Dev Ashish · · 3 min read

ask about a situation, not a definition. describe a duplicate payment, then reveal one stage at a time, the webhook arrived twice, the retries overlap, three servers handle the same event. engineers who have done this ask where the duplicate came from, reach for an idempotency key tied to the business operation, and worry about the window between checking and writing.

ask an engineer what idempotency means and almost everyone gets it right: doing something twice has the same effect as doing it once. that answer tells you nothing about whether they could stop a customer being charged twice.

so don’t ask for the definition. give them a situation and let it unfold.

the situation

reveal it one stage at a time. each stage depends on what they said before, so there is nothing to look up in advance.

stage one: the symptom

a merchant says a customer was charged twice for one order. both charges are real. where do you look first?

listen for whether they go looking for the source of the duplicate before proposing a fix. the useful places: the payment provider’s webhook logs, the order’s event history, whether the client retried, whether the user double-clicked. an engineer who jumps straight to “add a unique constraint” hasn’t asked what is unique yet.

stage two: the cause

the provider’s webhook for that payment arrived twice, 40 milliseconds apart. your handler checks whether the order is paid, then marks it paid and captures the charge. what happened?

this is the heart of it. the strong answer names the gap between checking and writing: both deliveries checked, both saw “not paid”, both went ahead. people who have been burned by this describe it without being prompted, and often add that webhooks are delivered at least once by design, so duplicates are normal, not a provider bug.

stage three: the fix that holds

make sure this can never happen again, even with three servers running the handler.

now you learn the most. listen for:

  • what the key is tied to. the provider’s event id, or the business operation (“capture payment for order 123”)? the second survives the provider sending two different events for the same thing.
  • atomicity. the check and the write must be one step, a unique constraint or an insert that fails on conflict, not a read followed by a write. a lock that lives in one server’s memory doesn’t help with three servers.
  • what a repeat returns. the first result, not an error, so a retrying client gets the answer it was waiting for.
  • how long keys are kept, and what happens to a retry that arrives after they expire.
  • the side effects outside your database, like the call to the payment provider itself, which needs its own idempotency key so a retry of your request doesn’t become a second charge there.

what the answers sound like

| what you hear | what it usually means | |---|---| | defines idempotency, then “use an idempotency key” and stops | has read about it | | asks where the duplicate came from before fixing it | has debugged it | | names the window between check and write unprompted | has been burned by it | | ties the key to the business operation, and makes the write atomic | has designed it | | mentions key expiry and the provider’s own idempotency | has run it for a while |

none of this needs a trick question, and an honest “i don’t know” at stage three is worth more than a confident guess. it tells you exactly where their experience ends, which is the point.

why this beats a coding test for this concept

a coding test checks whether someone can write a correct handler in isolation. the failure here isn’t in the handler. it is in the timing between two copies of it, which a single-threaded test never exercises. a conversation that unfolds shows whether they think in those terms on their own.

it also takes ten minutes, needs no setup, and works whether or not the candidate has public code.

the same pattern for other concepts

the shape works for most backend concepts: start from a symptom a real system would show, reveal the cause, then ask for a fix that survives scale. tail latency starts from “one customer’s calls are slow and the dashboard says everything is fine”. tenant isolation starts from “one customer saw another customer’s data for a second”. the definition is never the question.

this is what dwij calls a staged situation, and it is the core of a mapping. for why it matters more than ever now that every cv reads as a fit, see what is evidence-led hiring?

questions

what is idempotency in software?

an operation is idempotent when doing it twice has the same effect as doing it once. for a payment, it means a repeated request, a retried webhook or a double click charges the customer exactly once.

what is a good interview question for idempotency?

a situation that unfolds rather than a definition. for example, a customer was charged twice and the payment webhook arrived twice. ask where to look first, then reveal that the two events were processed 40 milliseconds apart, then ask how to make it impossible even with several servers.

what is an idempotency key?

a unique value attached to an operation, usually generated by the client or derived from the business event, so the server can recognise a repeat and return the first result instead of acting again. it has to be stored and checked atomically, or two copies can both get through.

what separates real experience from textbook knowledge here?

people who have handled duplicates in production ask where the duplicate came from, think about the time between checking and writing, choose what the key is tied to, and mention how long keys are kept. people who have only read about it define the term and stop at “use an idempotency key”.

hiring engineers? see the evidence before the interview.

join the pilot

read next

7 october 2026 · 4 min read

what is evidence-led hiring?

evidence-led hiring judges engineers on what they can show they understand and have done, with a source for every claim, instead of on how well a cv is written.