A merchant starts a catalog import while another merchant waits for an order update. Both operations use your integration, but they should not compete as one undifferentiated backlog. Adding workers can make this worse: more processes generate more requests without creating more upstream capacity.
This is an architecture explainer, not a new Salla announcement. Official documentation was checked on 30 September 2026. The queue design below is a proposed implementation pattern, not a built-in Salla feature.
Start with the upstream contract
Salla documents plan-dependent API limits and response headers, with a separate restriction for customer endpoints. Its response reference identifies HTTP 429 as rate limiting. Read the current contract and observed responses before configuring a client; do not copy a numeric limit into every worker indefinitely.
Keep endpoint-family rules separate from the general store budget. When documentation leaves reset units or quota scope unclear, verify actual responses and ask the provider before depending on an assumption. Treat absent or malformed headers conservatively.
Budget admission, not only execution
A recommended design gives every store a queue and a shared admission budget. A worker must obtain permission before sending a request. Multiple processes must coordinate this decision atomically; an in-memory counter in each process cannot enforce a shared limit.
Also cap concurrent requests. Throughput and concurrency describe different pressures: a slow response can occupy many sockets even when request frequency is low. Record the operation ID, store ID, endpoint family, attempt count, deadline and next eligible time with each job. Never put access tokens into job logs.
Keep one busy merchant from blocking the others
Consider a hypothetical integration with one large import and several small stores. A single first-in-first-out queue may put every urgent update behind the import. Instead, cycle through eligible stores and take a bounded amount of work from each. Within a store, reserve some capacity for time-sensitive operations while allowing background reconciliation to make progress.
Fair scheduling does not require equal throughput for every store. It requires an explicit allocation policy and a maximum acceptable waiting time. Do not let an unlimited high-priority stream starve reconciliation forever. Measure the age of the oldest eligible job, not only total queue length.
A 429 should reschedule work, not freeze a worker
When a response provides usable retry guidance, persist the next eligible time and release the worker. Do not hold a worker asleep for a long cooldown. If the affected budget is shared, other workers must see its cooldown too; otherwise they continue producing rejected requests.
Use bounded backoff with randomness when appropriate. AWS explains why jitter spreads synchronized retries. Apply that principle without retrying before a valid provider minimum. Put a deadline and attempt budget on the operation so an unavailable dependency cannot generate an immortal job.
A timeout leaves a write uncertain
A request that times out may already have changed remote state. Your local deduplication key prevents duplicate jobs inside your system, but it does not prove the upstream operation is idempotent. Before replaying a write, consult that endpoint's documented contract and look for a safe reconciliation path.
For example, if an operation has a documented external reference that can be queried, use it to check the result. If no reliable lookup exists, record an uncertain outcome and require a controlled recovery decision. Do not invent support for idempotency headers. This complements the existing guide to reliable webhook processing.
Make overload visible to the merchant
An accepted background job is not a completed synchronization. Show separate states for queued, running, delayed by provider limits, completed and needing attention. Give the merchant the last successful synchronization time and preserve progress across retries.
Operational dashboards should distinguish provider throttling from local saturation. Useful signals include oldest-job age per workload, completion latency, attempts per completed operation, uncertain writes and permanently failed jobs. Keep tenant identifiers out of unbounded metric labels; use access-controlled diagnostics for individual-store investigation.
Test the recovery path before increasing worker count
Simulate simultaneous workers for one store, missing headers, a prolonged cooldown, a timeout after a remote write and a large import beside small interactive workloads. Verify that the shared limiter is atomic, other stores continue progressing, deadlines terminate jobs and uncertain writes are not replayed blindly.
The release decision should be based on freshness and correctness, not raw requests per second. Start with a bounded rollout, review queue age and duplicate effects, then increase capacity only where your own system is the bottleneck. More workers are useful after admission, fairness and recovery are explicit—not as a substitute for them.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




