A cache is useful when it reuses an expensive result. It becomes dangerous when the application forgets that the source changes. Start by defining acceptable staleness. A category list may tolerate delay; a payment authorization or permission grant needs an authoritative check at the point of action.
Keep a clear source of truth
With cache-aside, the application checks Redis, reads the primary on a miss and stores a temporary copy. Database writes need an explicit invalidation strategy. Document affected keys beside the write path so that new features do not silently bypass it.
Treat keys as a data contract
A key such as products:1 ignores tenant, language, currency and authorization differences. Include dimensions that change the result. A version prefix also prevents old serialized shapes from being interpreted as a new format. Avoid personal data and secrets in key names.
catalog:v2:tenant:8f2:locale:ar:currency:SAR:page:1Tenant identity alone is insufficient when an administrator and employee see different fields. Separate the relevant permission scope or avoid caching the complete response. Test role boundaries within one tenant as well as boundaries between tenants.
Consider read-write races
A reader can fetch old data, another request can update the database and invalidate the key, and the original reader can then repopulate stale data. Deleting after a write does not eliminate every race. Depending on the harm, consider versioned values, coordinated refreshes or a short bounded stale window. Choose complexity from correctness requirements rather than a generic cache recipe.
Control expiry stampedes
When a popular key expires, many requests may hit the primary together. Coalesce concurrent refreshes, use carefully bounded locks where appropriate, and spread expiry with jitter. Locks need a timeout and a failure path; protecting a cache must not create indefinite waiting.
Measure useful behavior
Hit rate does not prove correctness or improved latency. Observe response time, value size, evictions and Redis failures. Decide how the application behaves without the cache. Unlimited fallback traffic during an outage can overload the primary database.
Start with a reversible rollout
Cache one expensive, low-risk read, record the freshness contract and test updates, concurrency and Redis unavailability. Keep business logic independent of the optimization. A cache should accelerate a understood operation, not hide unclear data ownership.
Scenario: changing a product price
A price may appear in product details, listings, search results and recommendation bundles. Deleting one key is insufficient when several representations contain copies. Map where the value is derived and choose invalidation for each representation. Checkout should calculate the amount according to authoritative current rules, rather than trusting a browser value or stale cached response, and communicate changes before confirmation.
Exercise a Redis outage
In a test environment, make Redis temporarily unavailable under a bounded request load. Observe database query volume and whether cache timeouts resolve quickly. Test recovery as well; returning cached values may no longer be appropriate. The goal is controlled behavior during failure and restoration, not merely an exception handler that hides an error message.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




