A small online furniture store sends its festive-sale email at 8 p.m. By 8:04, product pages take nine seconds to load. Nobody's code is broken. The database server is answering the same question ("what does the three-seater sofa cost, and is it in stock?") thousands of times a minute, and every answer costs it the same work as the first one did.
The fix usually isn't a bigger database. It's a place to keep answers you've already worked out, so you don't redo the work. For a large share of development teams, that place is Redis.
This guide explains Redis caching in plain language, then goes past the usual "put a cache in front of your database" advice. We'll look at what happens when the cache is missing data, when it disagrees with the database, when it runs out of memory, and when thousands of requests hit it at once. You don't need to be an engineer to follow along.
First, what Redis actually is
Redis is an in-memory database. That means it keeps its data in RAM, the short-term working memory of a computer, instead of on a hard drive or SSD. Reading from RAM is far faster than reading from disk, which is why Redis can usually answer a simple request in well under a millisecond.
Think of your main database as a filing cabinet in the basement: complete and reliable, but every trip downstairs costs time. Redis is the desk next to you, holding the few folders you open fifty times a day.
Redis stores data as keys and values. A key is a name, like product:4821:price, and the value is whatever you stored under that name. Values don't have to be plain text. Redis understands several data types, and a few of them matter a lot for caching:
▪ Strings hold a single value, such as a price, a chunk of JSON, or a piece of a rendered web page.
▪ Hashes hold several named fields under one key, a bit like a single row in a spreadsheet (a user's name, plan, and last login, for example).
▪ Sorted sets keep items ranked by a score, which is how leaderboards and "most popular this hour" lists are usually built.
▪ Any key can carry an expiry time, called a TTL (time to live). When the TTL runs out, Redis deletes the key by itself.
One more detail explains a lot of what comes later. Redis runs commands one at a time on a single main thread (versions 6.0 and newer use extra threads only for network traffic). That keeps Redis simple and very quick for small operations, but it also means one slow command makes every other request wait.
WHERE REDIS STANDS RIGHT NOW
A note on market size
Estimates for the in-memory database market disagree sharply. Mordor Intelligence (2026) puts the 2025 market at USD 7.08 billion, growing to USD 15.31 billion by 2031. SNS Insider (July 2026) values the same year at USD 17.53 billion. The two firms clearly define the category and measure it differently, so read either number as a signal of steady growth rather than a precise size.
A quick word on licensing before you commit
Business owners should run this past whoever handles legal questions. In March 2024, Redis Ltd. moved Redis from its permissive BSD license to source-available licenses (RSALv2 and SSPL). In response, the Linux Foundation backed a fork of the last BSD-licensed version, called Valkey, with support from several large cloud providers. Then in May 2025, Redis 8 added the AGPLv3, a license approved by the Open Source Initiative, as another option.
For teams running Redis as a cache behind their own app, little changes day to day. It matters more if you modify Redis and offer it to others as a service, since the AGPLv3 carries sharing obligations then. The patterns in this article work the same on Valkey.
The five main caching patterns, in plain words
Most caching strategies come down to two questions. Who puts data into the cache? And what happens when that data changes? Here are the five patterns you'll run into most, from the most common to the most specialized.
Cache-aside (lazy loading)
The app checks Redis first. If the answer is there (a "hit"), it uses it. If not (a "miss"), the app asks the database, stores the result in Redis with an expiry time, and returns it. The cache only fills with data somebody actually asked for.
It's the default choice because it's simple, and if Redis goes down, the app can still fall back to the database. The trade-off: the first request for anything is slow, and cached data can go stale until it expires or gets deleted.
Here's what it looks like in Python with the redis-py library. Even if you don't write code, the shape is easy to follow: check, fall back, store. (The db.fetch_product line stands in for your own database call.)
import json
import redis
r = redis.Redis(host="localhost", port=6379)
def get_product(product_id):
key = f"product:{product_id}"
cached = r.get(key)
if cached is not None:
return json.loads(cached) # hit: skip the database
product = db.fetch_product(product_id) # miss: do the slow work once
r.set(key, json.dumps(product), ex=300) # keep it for 5 minutes
return product
Read-through
This looks like cache-aside, except a caching library or framework does the fetching instead of your own code. Redis doesn't do this by itself. It keeps app code tidy but hides what's happening, which can make debugging slower.
Write-through
Every time the app saves something, it writes to the database and the cache in the same step, so freshly written data is always cached. The cost is slower writes and a cache full of data nobody may read. Adding a TTL lets unread data leave eventually.
Write-behind (write-back)
The app writes only to Redis, and a background process copies changes to the database later, often in batches. Writes feel instant, but this is the riskiest pattern here: if Redis crashes before the copy happens, those changes are gone. It suits view counters and analytics events. It doesn't suit payments, orders, or anything a customer would call support about.
Refresh-ahead
The cache renews popular items before they expire, such as refreshing a homepage list every 50 seconds when its TTL is 60. It works for a small set of predictable, heavily read items and wastes effort if you guess wrong about what's popular.
How the five patterns compare
★ PRO TIP
Start with cache-aside and a TTL on every key. Bring in other patterns only for the specific data that needs them. Using different patterns across one app is normal. Using two patterns for the same piece of data is where confusing bugs come from.
Deciding what goes in the cache, and for how long
Good candidates are read far more often than they change, take effort to produce, and do no harm if slightly old. Product details, category pages, user permissions, feature flags, and heavy report results all fit. A bank balance shown right before a money transfer does not.
Choosing a TTL is a business decision dressed up as a technical one. Ask, "If a customer saw this value five minutes late, would it matter?" A blog post can sit in cache for an hour; a flash-sale price might need ten seconds. Give nearly every key a TTL, since keys without one pile up until memory runs out.
Three habits that save trouble later
▪ Name keys with a clear pattern, such as app:tenant42:product:4821:v3. The version number at the end lets you change the stored format during a deploy without old and new code reading each other's data.
▪ Add a little randomness to TTLs. If 50,000 keys all get exactly 3,600 seconds at the same moment, they all expire together and the database takes the combined hit. A few minutes of random variation spreads that out.
▪ Keep values small. A few kilobytes is fine; several megabytes under one key slows every request that touches it.
Data gaps: when the cache has nothing, or half of something
A cache miss is normal. Some kinds of missing data, though, only cause trouble once real users arrive.
Asking for things that don't exist
Say a bot requests product IDs that were never created. Every request misses the cache, because there's nothing to cache, and hits the database. This is called cache penetration, and it can hurt as much as a traffic spike.
The simplest fix is negative caching. When the database says "not found," store that answer too, with a short TTL such as 30 to 60 seconds. For very large ID ranges, some teams add a Bloom filter, a compact structure that can say "this ID definitely doesn't exist" without asking the database. Redis 8 ships with Bloom filters built in; on other setups you may need a module.
Caching a partial answer
This gap is harder to spot. A product page pulls from the database and a reviews service. The reviews service times out, so the app builds the page without reviews and caches it. For the next ten minutes, every visitor sees zero reviews, long after the service recovered.
The rule: never cache a result built from a failure. Show that one user the reduced page but skip the cache write, or use a very short TTL. Some teams store a "complete: true" flag inside cached objects so readers can tell full results from partial ones.
The cold cache after a restart
When Redis restarts empty, every request misses at once. Redis can save to disk through RDB snapshots (a periodic copy) or AOF, an append-only log of writes replayed on startup, so a restarted server comes back with most of its data. For a pure cache, many teams skip that and instead preload the most-requested keys before sending traffic to a new server.
Conflicting signals: when the cache and the database disagree
The database is the source of truth. The cache is a copy. Every rule in this section follows from that.
Disagreement starts whenever data changes. A user updates their display name, and Redis still holds the old one. You can update the cached copy, or delete it and let the next read reload it. Deleting is usually safer, and the reason is a race condition, which is what programmers call two operations running at the same time in an order nobody planned for. Here's how updating the cache can go wrong when two people edit the same price:
Step 1 Request A writes a price of 500 to the database. Redis still holds the old price.
Step 2 Request B writes a price of 450 to the database a moment later.
Step 3 Request B updates the cache to 450.
Step 4 Request A, which was running slowly, now updates the cache to 500.
Result The database says 450, the cache says 500, and visitors see the wrong price until the TTL runs out. Had both requests simply deleted the key, the next read would have loaded 450.
Deleting doesn't solve everything. A reader can miss the cache and load the old value, then a writer updates the database and deletes the key, and then the reader, running late, writes the old value into the cache anyway.
Facebook engineers described this problem in their 2013 NSDI paper "Scaling Memcache at Facebook" and solved it with leases, tokens that become invalid if the key is deleted in the meantime. With Redis, you can get similar protection a few ways:
▪ Keep TTLs as a safety net. Even if a stale value slips in, it can't outlive its expiry.
▪ Store a version number or "last updated" time with each value, and refuse to overwrite a newer version with an older one. A short Lua script can do this check and the write as one atomic step, meaning no other command can run in between.
▪ Use a delayed second delete: delete the key on write, then again a second later to catch late stale writes. Crude, but common.
▪ Read from the primary database, not a replica, when refilling the cache right after a write.
Replica lag makes it worse
Many apps read from database replicas, copies that trail the main database by milliseconds to seconds. If a cache refill reads from a replica that hasn't caught up, it caches the old value and your invalidation did nothing. After a write, refill from the primary.
Signals across many app servers
If each app server also keeps a tiny local cache, deleting the Redis key won't clear those copies. Redis 6 added client-side caching for this, where Redis tells clients when a key they read has changed. Another route is announcing changes over Redis Pub/Sub, its built-in messaging feature.
Real-time decisions: what to trust the cache with
Much of the appeal of Redis caching is speed, and speed makes it tempting to push every live decision through Redis. Some real-time jobs fit it perfectly. Others need care.
Q. Can Redis handle rate limiting?
A. Yes. A common approach counts requests per user per minute with INCR and puts an expiry on the counter. Each command is atomic, so two requests can't both read "99" and both think they're the 100th.
Q. Should stock levels live in the cache?
A. Show them from the cache, but make the final call elsewhere. Displaying "3 left" from data that's a few seconds old is fine. Letting two customers buy the last item because both read a cached "1 in stock" is not. Deduct stock with an atomic operation (in the database inside a transaction, or in Redis with a single DECR or Lua script whose result you check), then reconcile with the database.
Q. What about sessions and logins?
A. Redis is a common session store because it's fast and handles expiry natively. If Redis loses data, users get logged out, so turn on persistence or replication if that's unacceptable. When a user changes their password or gets banned, delete their sessions immediately rather than waiting for a TTL.
How Redis behaves under pressure
Most caching problems never appear in testing. They appear at 8:04 p.m. on sale night, and knowing these patterns ahead of time is most of what separates good Redis performance from a bad evening.
The stampede
A popular key expires. In the next 200 milliseconds, 3,000 requests miss the cache together, and all 3,000 run the same expensive database query. This is a cache stampede, also called the thundering herd, and it can knock over a database that handles normal load with ease. Three fixes, often used together:
▪ A short lock. The first request to miss runs SET lock:product:4821 1 NX PX 5000 (NX means "only if this key doesn't exist," PX 5000 means "expire in five seconds"). Only the lock holder rebuilds the value; the rest wait briefly or serve the old copy.
▪ Serve stale while refreshing: keep returning the old value while one background job rebuilds it.
▪ Early random refresh. As expiry gets close, each request gets a small and rising chance of refreshing the value early. A 2015 VLDB paper by Vattani, Chierichetti and Lowenstein showed that this probabilistic approach spreads rebuilds out well.
Hot keys
Sometimes one key, like the homepage banner, draws a huge share of traffic. Adding Redis servers doesn't help, because a single key lives on a single server, which maxes out while the others sit idle.
Keep a one-second local copy of that key on each app server, or store identical copies under banner:home:1 through banner:home:8 and have each request pick one at random.
Big keys and slow commands
Because commands run one at a time, a command that takes 50 milliseconds holds up everything behind it for those 50 milliseconds. The usual culprits:
▪ KEYS, which scans every key and can freeze a large production server for seconds. SCAN does the same job in small steps.
▪ Reading or deleting a huge value, such as a list with a million items. UNLINK frees the memory in the background, where DEL blocks until it's done.
Redis keeps a log of slow commands, which you can read with SLOWLOG GET whenever latency jumps.
Running out of memory
RAM is limited and costs more than disk. You set a cap with the maxmemory setting, and the maxmemory-policy setting decides what happens when Redis reaches it. The default policy, noeviction, refuses new writes and returns errors. That's right for data you can't lose and wrong for a cache. These are the policies you'll choose between most often:
Eviction policies at a glance
One surprise: Redis's "least recently used" is an estimate. It samples a few keys (five by default) and evicts the best candidate among them, which is close enough in practice.
Watch fragmentation too. After many writes and deletes, Redis can hold more memory than its data needs. INFO memory reports this as mem_fragmentation_ratio, and values well above 1.5 deserve a look.
Scaling out with Redis Cluster
Redis Cluster splits data across servers using 16,384 hash slots. Commands touching several keys only work if those keys share a slot, which hash tags force: {user42}:cart and {user42}:profile land together because only the part in curly braces counts. Overuse this and you recreate the hot key problem.
Failover and the writes you can lose
Replicas copy the primary Redis server so one can take over if it fails. Replication is asynchronous by default, so the primary confirms a write before replicas have it. If the primary dies in that gap, those writes vanish. For a cache, the data just reloads. For sessions or write-behind data, it's a real loss. The WAIT command can hold a client until replicas confirm, at some cost to speed.
The network counts too
Redis may answer in a fraction of a millisecond, but each network round trip adds time, and a page making 40 separate Redis calls pays it 40 times. Pipelining (sending commands in one batch) and MGET, which fetches many keys at once, cut this sharply. Keep Redis in the same region as your app servers. A lot of Redis performance trouble turns out to be network distance rather than Redis itself.
Edge cases to write into your runbook
A runbook is the checklist your team follows when something breaks. These cases belong in it:
▪ A key that leaves out the customer or user ID. In an app with many client accounts, a key like dashboard:summary instead of dashboard:tenant42:summary can show one company's data to another. It's a one-word mistake with serious consequences, so check key names during code review.
▪ A personal page cached as a shared one. Cache a whole page that says "Hello, Priya," and the next visitor sees Priya's name. Cache the shared parts and add the personal parts on each request.
▪ Cached errors. Store an error message as a value and users see it until the TTL ends.
▪ Everything expiring at midnight. Setting every daily report to expire at exactly 00:00 schedules a stampede for yourself. Stagger the times.
Measuring Redis performance without fooling yourself
How the app feels on your laptop proves nothing. These numbers, mostly available through the INFO command, tell the real story.
▪ Hit ratio is hits divided by hits plus misses (keyspace_hits and keyspace_misses in INFO). There's no universal target. A low ratio often means TTLs are too short or keys are too specific, such as keys that include a timestamp.
▪ Latency percentiles matter more than averages. An average of 0.5 milliseconds can hide a 99th percentile of 40 milliseconds, which is what your unluckiest users feel. The redis-cli --latency tool helps here.
▪ A steadily climbing evicted_keys count means Redis doesn't have enough memory for what you're asking it to hold.
▪ Database load before and after matters most of all. Caching usually exists to protect the database, so track database CPU and queries per request next to your Redis performance numbers.
★ PRO TIP
Measure one slow page before you touch it: its response time at the 50th and 99th percentile, and how many database queries it makes. Add caching to that page only, then measure again. One clear before-and-after result convinces a team, or whoever approves the budget, far better than a general promise that caching makes things faster.
A rollout plan for a small team
If you're adding Redis to an existing app, this order lets you prove your caching strategies on a few pages before rolling them out widely:
1. Find the three slowest or most-requested read pages using your logs or monitoring tool.
2. Agree with the product owner on how many seconds of staleness each one can tolerate.
3. Add cache-aside with TTLs and versioned key names.
4. Add negative caching and a stampede lock for the busiest keys.
5. Set maxmemory and an eviction policy before launch, not after the first outage.
6. Create alerts for hit ratio, evictions, memory use, and latency.
7. Rehearse a Redis restart and a failover in testing, and watch the database.
When Redis isn't the right answer
Redis makes a sensible default, but it's one tool among several:
▪ Memcached is simpler and multi-threaded, and fine if you only need plain key-value caching.
▪ A CDN (content delivery network) caches public pages, images, and files on servers near your users, before traffic reaches you at all.
▪ Sometimes the real fix is a missing database index. Caching a slow query hides the problem, while adding the right index removes it.
Key takeaways
✓ Cache-aside with a TTL on every key covers most needs. Add other patterns on purpose, one data type at a time.
✓ When data changes, delete the cached copy instead of updating it, and keep TTLs as a safety net.
✓ Never cache results built during a failure, and cache "not found" answers briefly.
✓ Plan for stampedes, hot keys, and memory limits before the traffic shows up.
✓ The database holds the truth. The cache is a fast copy you can afford to lose.
Where to go from here
Back to the furniture store at 8:04 p.m. With cache-aside on product pages, slightly randomized TTLs, a lock on the busiest keys, and a sensible eviction policy, most sale traffic never reaches the database. Pages load from memory, and the database handles orders, the work that actually needs it.
That's the practical case for Redis caching. It moves repeated work out of the way so the important work has room. The harder part lies in the decisions around it: how stale is too stale, how you clean up after changes, and how the system acts when it's squeezed. Make those calls well for your first few pages, measure the difference, and let your wider caching strategies grow from real numbers instead of guesses.


