Caching definition
Caching is the practice of storing copies of data or computed results in a faster layer, such as memory, a CDN or the browser, so later requests can be served without repeating slow work. Good caching cuts latency and infrastructure cost dramatically, but every cache needs a strategy for keeping data fresh and invalidating it when the source changes.
How caching works
A cache sits between a consumer and a slower source. When a request arrives, the system checks the cache first. On a hit, it returns the stored copy in microseconds or milliseconds. On a miss, it fetches from the source, such as a database, an external API or a rendering step, returns the result and stores a copy for next time, usually with a time to live (TTL) after which the entry expires.
The benefit depends on the hit ratio, the share of requests answered from cache. Data that is read often and changes rarely, such as product catalogs, configuration or rendered pages, caches extremely well; data that changes on every request, such as a live account balance, may not be worth caching at all.
Layers of caching in a web application
Real systems stack several caches, each closer to the user than the last, and a miss at one layer simply falls through to the next. A single request for a product page might be answered by any one of these layers:
- Browser cache: controlled by Cache-Control and ETag headers, it avoids network requests entirely for static assets
- CDN cache: edge servers near users serve images, scripts and whole pages; see content delivery network
- Reverse proxy cache: Nginx, Varnish or a framework's built-in cache in front of the application
- Application cache: in-memory stores such as Redis or Memcached holding query results, sessions and computed values
- Database cache: buffer pools inside the database that keep hot data in memory
- AI caching: prompt caching and semantic caching of LLM responses to cut token cost and latency
Caching strategies and invalidation
Cache-aside, also called lazy loading, is the most common pattern: the application reads from the cache, falls back to the database on a miss and writes the result back. Write-through updates cache and database together, keeping them consistent at the cost of slower writes. Write-behind writes to the cache first and flushes to the database later, which is fast but risks losing data if the cache fails.
Invalidation is the hard part. TTLs are simple but leave a window of stale data. Event-based invalidation deletes or updates entries when the source changes, often triggered from a message queue. Versioned keys, such as content hashes in asset file names, sidestep invalidation for static files. Our Redis vs Memcached comparison covers the two most common in-memory caches.
Common caching pitfalls
A cache stampede happens when a popular entry expires and thousands of requests hit the database at once; request coalescing, locks or slightly randomized TTLs prevent it. Caching personalized responses at a CDN can leak one user's data to another, so private responses need Cache-Control: private or no-store. Caches also hide performance problems until they are cold after a deploy or restart, so load test with an empty cache too.
Nexzem designs caching layers as part of performance work on web and mobile backends, starting with measurements of where time is actually spent, because caching the wrong layer adds complexity and stale-data bugs without making users any faster.