A webhook is a fast signal, not an order ledger. If your Salla or Zid app sends orders to an ERP, warehouse or shipping provider, success means more than receiving the happy-path callback. The integration must tolerate duplicate deliveries, delayed retries and a period when no event is delivered at all.
What the official platform contracts actually say
Salla's webhook documentation lists events for order creation, updates, payment, cancellation, refunds and shipment changes. It documents signature or token-based verification, a connection/response window of about 30 seconds, and three attempts around five minutes apart when the endpoint does not return a successful response. Those retries improve delivery, but they do not make your database a replica of Salla.
Zid's webhook overview exposes order creation, status and payment-status events, alongside product and customer events. Conditions are currently documented for `order.create` and `order.status.update`, so an app can subscribe to a narrower stream where the use case permits it. The event names, payloads and conditions differ from Salla; a shared integration layer should normalize meaning after provider verification, not assume identical contracts.
Zid's circuit breaker changes the recovery question
Zid also documents webhook health tracking. An endpoint becomes degraded after 10 or more failures in a sliding hour and broken after 30 or more. A successful delivery resets the counter; HTTP 429 is documented as a live-but-rate-limited signal rather than a failure.
The important operational detail is what happens after the broken state: Zid says new dispatches are suspended and pending undelivered events are discarded. Recovery is explicit and requires a new target URL. Zid retries a failed event up to three times with delays of one, five and fifteen minutes before the exhausted delivery contributes to the health counter.
This is useful platform protection, but it means recovery cannot be “turn the worker back on and wait.” Your application needs to detect the gap and reconstruct state from authoritative APIs.
Put a durable inbox before business logic
Keep the public webhook handler small. Verify the provider signature or credentials against the raw request, identify the merchant, validate the minimum envelope, store the delivery in a durable inbox, then return 2xx. A queue worker can perform the slow work. Do not call an ERP, generate an AWB or send a campaign inside the request path.
Use a uniqueness key supplied by the provider when the contract exposes a stable event identifier. Otherwise build a carefully versioned fallback from provider, merchant, event type and an immutable source identifier; a payload hash can help detect repeats but should not be treated as a universal business identity. Preserve the raw payload with access controls and a retention policy, and store parsing errors separately from business-processing errors.
webhook -> verify -> durable inbox -> 2xx
|
v
queue -> normalized order projection
^ |
| v
reconciliation <- platform APIThe worker should upsert by `(platform, merchant_id, source_order_id)` and record the source update time when trustworthy. Each side effect needs its own idempotency boundary. Replaying an order update may refresh a local projection, but it must not create a second shipment or issue another refund.
Reconcile with the API, not with guesses
Run a scheduled reconciliation job per merchant. Query a bounded time window that overlaps the previous successful checkpoint, compare authoritative orders with the local projection, and enqueue repairs. The overlap deliberately rereads some records; idempotent upserts make that safe and protect against clock skew or late updates.
Salla's current List Orders endpoint documents sequential pagination, a maximum `per_page` of 30, a 15-minute cached pagination window and date filters. That contract shapes the backfill: partition long ranges into smaller windows, complete each page sequence, checkpoint only after the window succeeds, and respect rate-limit responses. Do not jump directly to a later page or rely on an old expanded response.
For Zid, the documented order list and webhook-health endpoints can support two different checks: business reconciliation asks whether order state matches, while health monitoring asks whether subscriptions are healthy, degraded or broken. Keep those alerts separate. A healthy endpoint can still contain a processing bug; a broken endpoint can leave the last successfully projected order looking perfectly valid.
Make provider differences explicit
Build a Salla adapter and a Zid adapter behind one internal contract such as `fetchChangedOrders`, `verifyWebhook`, `normalizeOrderEvent` and `inspectWebhookHealth`. Normalize only fields your product actually owns. Payment states, custom order statuses, branch inventory and refund semantics should retain the source value beside any internal mapping.
A canonical event should describe a fact, not invent one: `OrderPaymentStateObserved` is safer than `OrderPaid` when the platform payload does not prove final settlement. Version normalized payloads so a later mapping change can be replayed without rewriting history invisibly.
This provider-specific boundary complements the general reliable webhook processing guide and the rules in API contracts and error design. It also gives backend observability concrete signals: inbox age, duplicate rate, worker lag, reconciliation drift, provider 429s and broken-subscription count.
An incident playbook for missed orders
When the dashboard shows a gap, first stop destructive downstream actions, not ingestion. Confirm endpoint health and authentication, then record the last known successful event and reconciliation checkpoint. Repair the endpoint, recover or recreate subscriptions according to the platform contract, and backfill the affected time windows through the API.
Compare counts and source identifiers before reopening shipment or financial actions. Replay stored inbox records and reconciliation repairs through the same idempotent worker path. Do not patch the ERP manually while an automatic replay is still running unless every manual action is recorded as a deduplication key.
Finally, document the gap interval and the invariant that was checked: all source orders exist locally, every shipment side effect has one idempotency record, and no local order is newer than its confirmed source state without an explanation.
The practical decision
Use webhooks for low latency and APIs for recovery. Neither replaces the other. A durable inbox gives you evidence of what arrived; idempotent workers make replays safe; reconciliation finds what never arrived; provider-health monitoring tells you when the delivery channel itself has changed state.
The trade-off is extra storage, scheduled API traffic and more explicit state. For an app that only displays a non-critical notification, that may be excessive. For orders, payments, inventory or fulfillment, the cost is usually easier to operate than silent divergence between the merchant's store and the systems acting on its behalf.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




