Saudi commerce · OAuth security

Salla OAuth refresh tokens: rotate once, recover safely

A production design for Salla OAuth token refresh that prevents concurrent reuse, stores each rotation atomically and handles uncertain network outcomes safely.

TOPIC HUBE-commerce Engineering
Conceptual security illustration of one rotating OAuth token passing through a guarded single-flight gate into encrypted storage.
An editorial interpretation of the topic, followed by a practical execution diagram.

A Salla integration should treat token refresh as a single-writer state transition per store. Salla documents single-use refresh tokens: every successful exchange returns a new refresh token and invalidates the previous one. If two workers refresh the same merchant concurrently, reuse of the older token can revoke access and force the merchant to reinstall the app. The safe pattern is one bounded coordinator, an atomic token replacement and an explicit recovery state when the network outcome is unknown.

This is a practical Saudi-commerce engineering guide, not a claim about a new Salla release. The official documentation was checked on 5 October 2026. The locking, schema and rollout choices below are application patterns, not built-in Salla services.

Read the Salla contract before designing retries

Salla's authorization documentation describes OAuth 2.0 access through the merchant's consented scopes. It states that access tokens last two weeks, refresh tokens last one month, and a refresh token is single-use. A successful refresh produces a new pair and invalidates the previous refresh token. It also warns that reuse can revoke the associated access tokens and require app reinstallation.

That contract changes the meaning of a retry. Repeating a catalog read is usually harmless; repeating a refresh request is a security event. Generic HTTP middleware that retries POST after a timeout can damage authorization. Disable blind retries for the token endpoint and make refresh a named workflow with its own state, metrics and recovery policy.

The same page says published App Store apps use Easy Mode, where authorization events deliver tokens through the configured webhook; Custom Mode is for testing. Keep the implementation aligned with the current platform mode. Salla's security guidance also requires HTTPS and OAuth-authorized access for sensitive Merchant API data.

The race starts before expiry

Imagine ten queue workers processing one store. They read the same token row and see an access token close to expiry. Workers A and B both decide to refresh. A exchanges refresh token R7 and receives access token A8 plus refresh token R8. Before A commits R8, B sends R7. Because R7 has already been consumed, the provider sees reuse rather than a normal second request.

A transaction around each local update does not solve that sequence if the remote call happens outside a shared coordination boundary. Last-write-wins storage is unsafe too: B cannot overwrite A with an error, and A's successful response may arrive after another worker has marked the connection for reauthorization. The operation needs one owner per merchant and a monotonically increasing token version.

Refresh proactively before access expiry, but add deterministic jitter per store so thousands of merchants do not refresh in the same minute. A 401 may have several causes; route it through the same coordinator, allow at most one evidence-based refresh attempt, then surface the actual authorization failure.

Store a versioned secret, not a loose pair of strings

Keep one row per installed merchant and app. Useful fields are merchant_id, app_id, encrypted_access_token, encrypted_refresh_token, access_expires_at, refresh_expires_at, token_version, status, refresh_lease_owner, refresh_lease_until and updated_at. A unique constraint on merchant_id plus app_id prevents two installation rows from becoming competing authorities.

Encrypt both tokens with an envelope key managed outside the database. Restrict decryption to the refresh and API-client services. Never put credentials in queues, URLs, analytics events or exception messages. Logs can carry merchant ID, token version, correlation ID, outcome class and latency without containing the token or its hash. Request only the Salla scopes the product requires.

Every worker carries the token_version it read. After waiting for another refresh, it reloads the row: if the version advanced and status is active, it uses the new access token instead of refreshing again. Installation and app-update webhooks should enter through the same atomic writer so an authorization event cannot race a scheduled refresh.

One coordinator per store acquires a lease, refreshes once, atomically replaces both tokens and releases waiters onto the new version.
One coordinator per store acquires a lease, refreshes once, atomically replaces both tokens and releases waiters onto the new version. Open for a larger view

Coordinate one refresh per merchant

A durable lease works across processes and survives a worker crash. Acquire it with one conditional statement: only a row whose lease is absent or expired may change to the current worker ID and a short deadline. If acquisition fails, do not call Salla. Wait briefly, reload the row and observe whether another worker advanced token_version or marked reauthorization.

The coordinator reads the encrypted refresh token and expected version, calls the documented token endpoint once, then commits both returned tokens in a short transaction. The update compares merchant ID and expected token_version. If that comparison fails, do not use the response normally: another authority changed the row while the request was in flight.

const lease = await acquireRefreshLease(merchantId, workerId, 30_000);
if (!lease.acquired) return await waitForNewerToken(merchantId, lease.version);

try {
  const before = await loadEncryptedTokenState(merchantId);
  const rotated = await sallaRefreshOnce(decrypt(before.refreshToken));
  return await replaceTokensAtomically({
    merchantId,
    expectedVersion: before.version,
    accessToken: encrypt(rotated.access_token),
    refreshToken: encrypt(rotated.refresh_token),
    nextVersion: before.version + 1
  });
} catch (error) {
  await classifyRefreshFailure(merchantId, error);
  throw error;
} finally {
  await releaseRefreshLease(merchantId, workerId);
}

This is an architectural sketch, not a drop-in SDK. Use the exact response fields and expiry units documented and observed for the selected Salla flow. Validate lease ownership before the final write. Bound network timeout below the lease duration, renew only under explicit ownership, and make release conditional on worker ID so a late process cannot clear its successor's lease.

PostgreSQL also provides application-defined advisory locks. They can serialize a store key, but session locks depend on connection ownership and transaction locks end with the transaction. A lease row is often clearer with poolers and remote calls. If advisory locks are used, pin the session, release in finally and monitor pg_locks.

A lost response is an uncertain security state

The hardest case is a timeout after Salla accepted R7 but before the client received R8. Retrying R7 can trigger reuse detection, while the client cannot reconstruct R8. Local locking cannot repair a response it never received. Mark the connection refresh_unknown, stop further refresh attempts and prevent destructive downstream actions until authorization is recovered through a documented path.

Do not classify every timeout as invalid_grant, and do not immediately ask the merchant to reinstall before checking whether another coordinator committed a newer version. Reload the row, inspect correlation and version, and allow a short bounded interval for a late successful writer. If no newer token exists, present a precise reconnect action and preserve pending work for replay after authorization returns.

RFC 9700 explains the trade-off behind refresh-token rotation: replay detection cannot know whether the legitimate client or an attacker submitted the invalidated token, so revoking the active token stops the attack at the cost of a new authorization grant. That behavior is protection, not a transient error to defeat with another retry.

Separate authorization recovery from business retries

When a worker cannot obtain an access token, keep the business job durable but blocked by authorization. Do not consume its retry budget every minute. Use connection states such as active, refreshing, refresh_unknown, reauthorization_required and revoked, and let job admission check that state before contacting Salla.

After the merchant reconnects, save the new pair as a fresh version, clear the block and resume jobs through their existing idempotency keys. Order, shipment and payment writes still need their own uncertainty handling; a valid token does not make a remote mutation idempotent. The Salla and Zid webhook recovery guide covers the separate path for delayed or missed events.

Avoid a global refresh queue that lets one broken merchant block the rest. Serialization belongs to the merchant connection, while queue capacity and alerting operate across the fleet. Set a maximum wait so a slow token endpoint cannot hold request threads indefinitely.

Observe coordination without leaking credentials

Measure refresh attempts, successful rotations, lease contention, wait duration, token-version advances, invalid_grant responses, uncertain outcomes and merchants awaiting reauthorization. Break metrics down by app version and outcome class, but do not use unbounded merchant IDs as labels. Use access-controlled traces for individual investigations.

Concurrent refresh attempts for the same version indicate a coordination bug even when one succeeds. Repeated lease expiry suggests a timeout or worker crash. A spike in reauthorization_required after a deployment should stop the rollout. Track the age of blocked business jobs so authorization failures do not become silent order drift.

An audit record can say which service requested refresh, the expected version, lease result, provider response class, committed version and recovery state. It must never contain access tokens, refresh tokens, client secrets or full authorization headers.

Test failure ordering before production

Test two simultaneous workers, a crash after lease acquisition, timeout before the provider reads the request, timeout after token consumption, late response after lease expiry, an Easy Mode authorization webhook during refresh and merchant uninstall. Verify that each refresh token is presented once and an older version cannot overwrite a newer one.

Then test a request that starts before rotation, a 401 during rotation, a queued order after reauthorization and a deployment with mixed application versions. Use a demo store and synthetic data as Salla recommends; do not test destructive recovery against live merchant orders.

This design is valuable for multi-store apps with multiple workers, scheduled jobs or bursty webhooks. A single-process prototype may begin with an in-memory mutex, but it becomes unsafe as soon as another process, region or runner can refresh the same merchant. The production invariant is clear: one refresh token is presented once, its replacement is stored atomically, every waiter observes the new version, and uncertainty stops automation instead of gambling with the merchant's connection. For implementation support, see Salla and Zid integration engineering.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceSalla & Zid apps and merchant toolsRelevant projectMember Plus