“Add Redis” is often a diagnosis failure
When latency rises, caching is attractive because it can reduce database work quickly. But if the dominant problem is an inefficient query, missing index, lock contention or an oversized synchronous transaction, a cache can hide the symptom while increasing invalidation and consistency complexity.
First classify the bottleneck
| Signal | First investigation | Redis likely? |
|---|---|---|
| Repeated identical read-heavy queries | Query frequency + freshness tolerance | Often |
| Slow unique queries | Execution plan, indexes, data model | Usually not first |
| Write contention | Transaction scope, locks, hot rows | Rarely the core fix |
| Expensive computed result reused broadly | Recompute cost + invalidation model | Often |
| Session/rate-limit ephemeral state | Consistency and expiry needs | Good fit |
The cache tax
A cache introduces key design, invalidation, TTL policy, stampede handling, memory pressure, observability and degraded-mode behavior. Count that operational cost in the decision matrix instead of treating Redis as “free speed”.
Decision sequence
- Measure the hot path and identify what consumes time.
- Fix obviously inefficient database work first.
- Estimate the theoretical gain from caching the repeated work.
- Define freshness and invalidation requirements.
- Load-test both the happy path and cache-miss/stampede path.
Calculate the ceiling before adding infrastructure
Synthetic arithmetic, not a benchmark: assume an average request takes 200 ms, of which 120 ms is eligible repeated database work. At an assumed 80% hit rate and 5 ms cache lookup on every request, the simplified average becomes 80 + 5 + (0.20 × 120) = 109 ms. This excludes network variation, serialization, invalidation, misses under load and other costs.
If eligible work is only 20 ms, the same assumptions produce 180 + 5 + (0.20 × 20) = 189 ms. The theoretical gain is now small. Neither calculation predicts p95 or p99: averages and tail latency are different quantities.
| Before deciding | Record |
|---|---|
| Database contribution | Measured request spans, plans, row estimates and actual work |
| Freshness | Maximum tolerated stale data for this operation |
| Cache failure | Load on the database during cold start and cache outage |
| Exit condition | Target latency and error rate under a reproducible load |
A useful next experiment
Measure one representative read path on a safe test workload. Inspect its query plan, correct the highest-impact inefficiency, then repeat the same workload. Only compare the cache option after defining hit-rate and freshness assumptions. Keep the original dataset, workload shape and configuration with the result.
Source and interpretation
PostgreSQL's EXPLAIN documentation explains plans and estimates. EXPLAIN ANALYZE executes the statement; use a suitable test environment. The illustrative latency arithmetic above is SYSLUME's model, not a PostgreSQL or Redis performance claim.