E-commerce engineering · Shopify

Reliable Shopify webhooks: durable inbox and reconciliation

A production design for authenticating, deduplicating, ordering and recovering Shopify webhook-driven order and inventory synchronization.

TOPIC HUBE-commerce Engineering
Conceptual architecture illustration of commerce events passing through a guarded durable inbox, worker and reconciliation loop.
An editorial interpretation of the topic, followed by a practical execution diagram.

A reliable Shopify webhook integration does not treat a successful HTTP request as a complete business event. It verifies the raw-body HMAC, records every delivery in a durable inbox before replying, deduplicates by Shopify's delivery identifier, processes side effects asynchronously and runs reconciliation against the Admin API. This design matters for apps that synchronize orders, inventory, fulfillment or customer state, because Shopify documents that duplicate delivery can occur, ordering is not guaranteed and webhook delivery itself is not guaranteed.

This is a practical architecture guide based on Shopify's official developer documentation, verified on 5 October 2026. It is not an announcement of a new Shopify feature. The inbox schema, worker topology, idempotency keys and recovery controls below are application-design recommendations; Shopify's documentation defines the delivery contract and headers.

Start with the delivery contract

Shopify webhooks are near-real-time notifications that reduce repeated polling. An app subscribes to topics and receives HTTPS POST requests when matching changes happen. That makes webhooks useful triggers, but not an ordered change-data-capture log. Shopify explicitly says ordering is not guaranteed within one topic or across topics, and that an app should not rely only on webhooks because delivery is not guaranteed.

The transport budget is deliberately tight. Shopify documents a one-second connection timeout and a five-second total request timeout for HTTPS delivery. Any response outside the 200 range is considered a failure. The practical consequence is that the receiver should perform only the minimum synchronous work: read the unmodified request body, verify authenticity, insert one durable inbox record and return 200. Parsing a large graph, calling another API or updating five business tables before acknowledgement consumes the wrong reliability budget.

Shopify can retry a failed HTTPS delivery up to eight times over four hours. For subscriptions created through the Admin API, repeated failures can remove the subscription. App-specific subscriptions declared in app configuration have different lifecycle behavior and are not automatically deleted for the same failure condition. Monitor the exact subscription type you operate rather than assuming a retry queue protects it forever.

Verify the raw request before trusting it

Shopify signs HTTPS webhook requests with an HMAC in `X-Shopify-Hmac-SHA256`. Verification must use the raw request body and the app's client secret. If a framework parses JSON, normalizes whitespace or changes encoding before verification, the calculated digest may differ even when the payload is legitimate. Middleware order is therefore part of the security boundary.

Compare digests with a timing-safe operation and reject invalid requests before using the shop domain, topic or payload fields. Keep the secret in a managed secret store and support deliberate rotation. Do not log the raw body indiscriminately: order and customer payloads can contain personal data. A useful security log records the request ID, topic, shop identity after validation, verification result and a bounded payload hash.

const raw = await readRawBody(request);
const supplied = request.headers.get('x-shopify-hmac-sha256');
if (!verifyShopifyHmac(raw, supplied, activeSecrets)) {
  return new Response('invalid signature', { status: 401 });
}

await inbox.insertIfAbsent({
  shop: request.headers.get('x-shopify-shop-domain'),
  topic: request.headers.get('x-shopify-topic'),
  deliveryId: request.headers.get('x-shopify-webhook-id'),
  eventId: request.headers.get('x-shopify-event-id'),
  triggeredAt: request.headers.get('x-shopify-triggered-at'),
  rawBody: encrypt(raw),
  receivedAt: now()
});
return new Response(null, { status: 200 });

The sketch is provider-neutral application code, not a copy of a Shopify SDK. Use the official library and runtime guidance for the framework you deploy, and test against the exact byte sequence delivered to production.

Build a durable inbox before the queue

A queue is not automatically an inbox. If the receiver acknowledges before publishing to the queue, a crash between those operations loses the event. If it publishes first and then fails before acknowledgement, Shopify may retry and create a duplicate. The safe boundary is one durable insert with a uniqueness constraint on `(shop_id, delivery_id)`, followed by 200. A separate dispatcher can publish committed rows to the work queue.

Store the delivery ID, event ID when present, topic, validated shop, API version, triggered time, received time, encrypted raw payload or a compliant reference, processing state, attempt count and last error. Keep payload retention proportional to replay and audit needs; do not turn the inbox into an indefinite copy of customer data. A transactional outbox is useful when the inbox database and worker broker are different systems.

Shopify's documentation distinguishes two identifiers. `X-Shopify-Webhook-Id` identifies a delivery and is appropriate for duplicate suppression. `X-Shopify-Event-Id` can correlate multiple webhook messages generated by the same merchant action. Do not collapse them into one column. Two distinct subscriptions can legitimately produce related deliveries with one event ID, while the same delivery ID should not apply its effects twice.

Verify the raw request, persist each delivery before acknowledgement, apply idempotent effects asynchronously and repair missed changes through reconciliation.
Verify the raw request, persist each delivery before acknowledgement, apply idempotent effects asynchronously and repair missed changes through reconciliation. Open for a larger view

Make every side effect idempotent

Deduplicating the HTTP delivery is necessary but insufficient. A worker can finish an external API call and crash before marking the inbox row complete. The next attempt then repeats the call. Put an idempotency boundary around each business effect, not only around receipt.

For an order projection, use a key such as `(shop_id, source_order_id, source_updated_at, projection_version)`. For an outbound fulfillment action, persist an operation row with a stable business key before calling the carrier or ERP. If that provider supports an idempotency key, send the same key on every retry. If it does not, query for existing state before creating again and route ambiguous outcomes to review.

Keep handler stages explicit: validate schema, load tenant policy, acquire the resource version, calculate the intended transition, write local state, perform guarded external effects and record completion. A retry should resume or safely replay a stage. “Exactly once” is not a transport promise; it is an application invariant assembled from uniqueness, transactions and idempotent effects.

Protect state from out-of-order events

Arrival order is not business order. An `orders/updated` delivery may reach the app before an earlier `orders/create`, or two updates can cross in flight. Shopify recommends using `X-Shopify-Triggered-At` or resource timestamps such as `updated_at` to understand when the event occurred. Treat those values as comparison evidence, not as permission to overwrite blindly.

Maintain a per-resource source version or latest accepted `updated_at`. Apply an event only when it advances the projection, or fetch authoritative current state when the transition is ambiguous. A stale event can still be useful for audit, but it should not move a paid order back to an earlier state. If two topics contribute to one aggregate, define a merge rule rather than relying on a global queue order that Shopify does not provide.

Partitioning a worker stream by shop and resource ID can reduce races, but it cannot reconstruct information that was never delivered. It also creates hot partitions for large stores. Use bounded ordering where it protects a clear invariant, and let independent resources process concurrently.

Reconcile because delivery is not guaranteed

The recovery loop is not an emergency script; it is part of the normal design. Shopify recommends reconciliation jobs when an app cannot afford to miss changes, using Admin API filters such as `updated_at` to retrieve data changed since the last successful job. Maintain a per-shop checkpoint and query an overlap window, because clocks, pagination and transaction visibility can create boundary gaps.

For example, a job running every ten minutes can query from `last_checkpoint - overlap`, page until the current safe horizon, and feed every result through the same idempotent projection path used by webhooks. Advance the checkpoint only after all pages are durably processed. The overlap deliberately creates duplicates, which the system already knows how to absorb.

Do not reconcile the entire catalog for every store on every run. Use resource-specific cursors, rate-limit budgets and priority by business risk. Orders and fulfillment may need a tighter recovery objective than product copy. Run a periodic inventory of active subscriptions as well, especially for shop-specific Admin API subscriptions that can be removed after repeated delivery failure.

Choose and version subscriptions deliberately

Shopify recommends app-specific subscriptions in `shopify.app.toml` for events required across every installation. They are deployed with the app configuration and provide one source-controlled contract. Shop-specific subscriptions created through the GraphQL Admin API fit per-merchant or runtime-specific needs. Record why each subscription exists, its owner and expected topic set.

Use subscription filters to reduce irrelevant deliveries and `include_fields` when a smaller payload is genuinely enough. These controls reduce network and parsing work, but they are also data contracts. A filter that excludes a transition or an omitted field needed by recovery can silently break the projection. Test sample resources at the boundary and keep a contract test for each configured filter.

Webhook payloads are versioned by Shopify API version. Pin the intended version, inspect the delivered version header and rehearse upgrades against captured compliant fixtures. Update parsers additively first, deploy them before changing the subscription version, and retain a short compatibility window. A webhook endpoint that accepts only one exact schema makes platform-version migration needlessly risky.

Monitor the pipeline, not only HTTP errors

Shopify's troubleshooting guidance emphasizes delivery count, failure rate and response time. Add application signals that expose silent drift: HMAC rejection rate, inbox insert latency, duplicate-delivery rate, oldest unprocessed inbox age, worker attempts, dead-letter count, event-to-projection lag, reconciliation discoveries and subscription inventory mismatch.

Measure p90 and p99 acknowledgement latency below the five-second ceiling; an average can hide a tail already causing retries. Alert separately when the receiver is healthy but the worker backlog grows. A 200 response proves durable acceptance only if the insert boundary is correct; it does not prove the order reached the ERP.

Trace one event using internal IDs rather than customer details. Useful fields are delivery ID, event ID, shop ID, topic, resource ID, source timestamp, inbox record, worker attempt and effect key. Limit metric labels: shop IDs and order IDs belong in access-controlled logs or traces, not unbounded time-series dimensions.

Know when webhooks are not enough

Use webhooks for low-latency triggers and incremental synchronization. Use reconciliation for completeness. Use a bulk export or paginated API job for initial backfill, schema migration or a full audit. Polling can be simpler for a small integration with loose freshness needs and few stores, although it should still respect API limits and checkpoints.

Do not use a webhook body as permanent evidence of current Shopify state when the workflow requires the latest authoritative resource. Fetch current state after coalescing bursts if the decision depends on a complete object. Conversely, do not issue an Admin API read for every event when the payload already contains the stable fields needed; that creates avoidable latency and rate-limit pressure.

The production rule is straightforward: authenticate bytes before trust, store before acknowledge, distinguish deliveries from events, make effects idempotent, compare source versions instead of arrival order, and repair gaps through reconciliation. That architecture turns Shopify webhooks from a fragile callback into a dependable synchronization signal. For broader patterns, see designing reliable webhook processing, the Salla and Zid webhook recovery guide and backend API integration services.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceBackend engineering & API integrationsRelevant projectMember Plus