Integrations · Webhooks

Designing webhooks for duplicates, retries and delayed events

A durable inbox, signature checks and idempotent processing for payment and commerce integrations.

TOPIC HUBAPIs, SaaS & System Architecture
Message paths converging on one verified outcome, illustrating webhook processing.
An editorial interpretation of the topic, followed by a practical execution diagram.

A webhook is a delivery attempt, not a promise of exactly-once, ordered processing. If every payment notification extends a membership, a retry may grant the same period twice. Separate reliable receipt from business execution and assume duplicates, delayed events and interrupted workers will occur.

Verify the provider’s contract

Use the documented signature mechanism. Providers such as Stripe verify the raw request body; parsing and reserializing JSON can change the bytes. Keep secrets outside source code, bound request size and reject invalid signatures. An event name is not authentication.

Persist a durable inbox

After verification, store the event with its provider, external account, event identifier and processing state. A unique key such as provider + account_id + event_id can deduplicate deliveries within the appropriate scope. Acknowledging before durable storage can lose an event if the process dies before queue submission. A polling inbox worker or a transactional outbox can close that handoff gap.

Verify signature -> persist event with unique key -> acknowledge
Worker -> claim event -> apply business transition -> mark processed
Failure -> record reason -> bounded retry or manual review

Deduplicate business effects as well

Different event identifiers may describe the same commercial fact. For a membership renewal, associate the granted period with the payment or billing cycle and enforce uniqueness there. When both records live in one database, update the membership and effect ledger in a transaction. An external action needs provider-supported idempotency or reconciliation; a SQL transaction cannot atomically include an arbitrary remote request.

A durable inbox precedes acknowledgement; the worker protects the business effect.
A durable inbox precedes acknowledgement; the worker protects the business effect. Open for a larger view

Expect delayed state transitions

A cancellation can arrive before an older subscription update. Define permitted state transitions instead of treating delivery order as business order. If the provider supports it, retrieve current state when events conflict. Retain enough event context for investigation under an appropriate retention policy, without copying sensitive payloads into unrestricted logs.

Exercise the crash boundaries

Replay the same event, deliver two copies concurrently, and stop the worker after the business change but before completion is recorded. Test unavailable storage during receipt and a failed email after successful payment. Each restart should preserve one business effect and expose unfinished work.

Monitor backlog age

Track the oldest unprocessed event, retry count and failure rate, not only queue length. Reconcile important entities periodically against the external source. Providers have different retry and ordering contracts, so document each integration rather than assuming Stripe’s behavior applies universally.

Scenario: a successful payment and a crashed worker

Imagine a worker granting membership and crashing before marking the inbox event complete. On restart, the event still looks unfinished. The business-effect ledger should show that the payment was already consumed, allowing completion without granting another period. When membership and effect records share a database, update them in one transaction. Notifications can have their own delivery ledger so a later retry does not necessarily send another message.

What should operators see?

Show event and operation identifiers, state, attempt count, the latest failure and the associated business record. A replay action should use the same idempotency protections as ordinary processing. Distinguish a temporary failure from an unknown remote outcome, which may need reconciliation before another attempt. Recovery should not depend on copying commands manually out of production logs.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceBackend engineering & API integrationsRelevant projectWhatsApp Hero