A payment succeeds, but the booking screen still says “pending.” Someone retries the workflow. Now the customer has two confirmation emails, or a second record appears in the accounting system.
The difficult part is deciding whether the first attempt failed before doing any work, or completed its work before the response was lost. A retry policy needs to answer that question before repeating an action.
This guide describes a design review and test plan. The examples are hypothetical; they are not customer incident reports or measured performance results.
1. Separate delivery from the business action
An incoming payment event and a confirmed booking are different records. Keep an event receipt with its provider event identifier, received time, processing status and related business record. Then make the business transition explicit: for example, change a reservation from awaiting payment to confirmed only when its current state allows that transition.
Stripe’s webhook documentation explains that events can be delivered more than once and are not guaranteed to arrive in order. A second delivery should therefore be an expected input, not an exceptional incident.
2. Decide what identifies the operation
Two deliveries of the same event can share an event identifier. Two different events may still refer to the same payment or business operation. Define both levels: a receipt identifier for the incoming event and a stable operation identifier for the action you intend to perform.
For a hypothetical booking, the operation might be “confirm booking B-104 for payment P-208.” Repeating that operation should not create another booking or resend its notification automatically.
A check followed by a write is not enough when two workers can run together. Use a database uniqueness constraint, transaction, atomic claim or an equivalent mechanism supported by your stack. Document what happens when the second worker cannot claim the operation.
3. Treat an uncertain result differently from a rejection
A validation error usually calls for correcting the input. A temporary service failure may permit a bounded retry. A timeout after sending a request is ambiguous: the remote service might have completed the action.
Before retrying an ambiguous write, look up its result using the operation identifier or use the provider’s supported idempotency mechanism. Stripe documents idempotency keys for supported API requests. Those keys do not replace your application’s event receipt and business-state controls.
4. Bound the retries and give someone ownership
Define the maximum attempts, delay, final failure state and person responsible for recovery. Include a route for reviewing uncertain outcomes instead of endlessly retrying them.
The recovery view should connect the event, operation and business record. Record enough information to explain the decision without copying secrets or unnecessary customer details into logs.
5. Test the awkward sequences before launch
Use synthetic records in a test environment and verify the final business state, not just whether a workflow reports success:
- Deliver the same event twice, then twice at the same time.
- Deliver a payment confirmation after a reservation-expiry job.
- Simulate a remote write succeeding while its response is lost.
- Stop a worker after the business update but before recording completion.
- Replay an already completed operation through the recovery interface.
For each case, record the expected state, actual state and number of side effects. A successful recovery should not silently create a second charge, booking or notification.
Apply the review to a marketplace
In our HomeSchool Matchmaker build story, booking reservations, marketplace payments and recurring memberships have separate lifecycles. That is useful context for deciding which operation a retry belongs to. The case study describes the architecture; it does not claim that the test plan above was performed on the client app.
Before adding another retry, write down the action, its stable identifier, the allowed state transition and the recovery owner. Those four details make an automation easier to operate when delivery becomes uncertain.