Caching in Java: Caffeine In-Process, Redis Distributed

Last Diwali, our checkout service got its first real flash sale. At noon the marketing email went out; by 12:04 the p99 checkout latency had climbed from 120 ms to 8 seconds, the database CPU was pegged at 100%, and the connection pool was exhausted — new checkouts were timing out waiting for a connection. The queries themselves were fine (we had fixed the N+1 problems covered in this track's companion post, "JPA Performance: LAZY/EAGER, N+1, Fetch Joins & Locking"). The problem was arithmetic: 40,000 checkouts a minute, and every single one ran the same indexed price lookup against the database. The database was answering the same question 40,000 times a minute, correctly, and dying.

So we did what every team does first: we cached everything, aggressively, with a long TTL. The database recovered overnight. Then came the second incident. A price drop scheduled for 10:00 AM didn't reach customers until the cache entries expired hours later — shoppers were charged the old, higher price, and support spent a week issuing refunds. The first outage was a missing cache. The second was a cache nobody had thought about invalidating. This post is about the deliberate middle: cache-aside with Caffeine inside one JVM, Redis when the cache must be shared, and the invalidation and stampede discipline that keeps both honest.

Cache-aside: the only pattern you need first

There are several cache patterns (read-through, write-through, write-behind), but cache-aside is the one to learn first because the application stays in charge: your code checks the cache, falls back to the database on a miss, and stores the result. The cache never invents data and never talks to the database on its own:

Checkout code price = cache .get(id, dao::find) Caffeine cache hot keys live here TTL 10 min · stats on Product DB source of truth slow, shared, billed 1. lookup 2a. HIT — return now DB never sees the request 2b. MISS 3. load, store, then return The cache never talks to the DB on its own. Your code decides when data enters. That is the whole of cache-aside: the application owns both the read path and the write path.

For the in-process cache we use Caffeine (com.github.ben-manes.caffeine:caffeine:3.1.8). It is the modern successor to Guava's cache: lock-striped reads, a W-TinyLFU eviction policy with near-optimal hit rates, and an API that reads like the diagram above. Here is the worked demo — a product catalog the checkout reads, a fake DAO standing in for the database (the connection and pool mechanics behind a real DAO are this track's "JDBC First: Connections, Pools & Transactions" post), and a Caffeine cache in front of it:

package com.javamakeuse.catalog;

/** Product from the checkout-service catalog. priceCents avoids float rounding. */
public record Product(long id, String name, long priceCents) { }
package com.javamakeuse.catalog;

import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.atomic.AtomicLong;

/**
 * A stand-in for the real product DAO. It behaves like a database:
 * slow (a few ms per call) and it counts every call so the demo
 * can prove how many "DB hits" the cache absorbed.
 */
public class FakeProductDao {

    private final Map<Long, Product> db = new ConcurrentHashMap<>();
    private final AtomicLong dbHits = new AtomicLong();

    public FakeProductDao(int products) {
        for (long i = 1; i <= products; i++) {
            db.put(i, new Product(i, "SKU-" + i, 999 + i * 10));
        }
    }

    public Product findById(long id) {
        dbHits.incrementAndGet();
        try {
            Thread.sleep(2); // simulate a cheap indexed SELECT
        } catch (InterruptedException e) {
            Thread.currentThread().interrupt();
        }
        return db.get(id);
    }

    public long dbHits() {
        return dbHits.get();
    }
}
package com.javamakeuse.catalog;

import com.github.benmanes.caffeine.cache.Cache;
import com.github.benmanes.caffeine.cache.Caffeine;
import com.github.benmanes.caffeine.cache.stats.CacheStats;
import java.time.Duration;

/**
 * Cache-aside with Caffeine, against the catalog the checkout reads.
 *
 * Pattern: on a request, ask the cache first; only on a miss does the
 * loader touch the database, and its result is stored for next time.
 */
public class CacheAsideDemo {

    public static void main(String[] args) {
        FakeProductDao dao = new FakeProductDao(200);

        Cache<Long, Product> products = Caffeine.newBuilder()
                .maximumSize(10_000)
                .expireAfterWrite(Duration.ofMinutes(10))
                .recordStats()
                .build();

        // Phase 1 — cold cache: 20 distinct products, every one a miss.
        for (long id = 1; id <= 20; id++) {
            long key = id;
            products.get(key, k -> dao.findById(k));
        }

        // Phase 2 — warm cache: 500 reads over the same 20 keys, all hits.
        for (int i = 0; i < 500; i++) {
            long key = (i % 20) + 1;
            products.get(key, k -> dao.findById(k));
        }

        CacheStats s = products.stats();
        System.out.println("requests      = " + s.requestCount());
        System.out.println("hits          = " + s.hitCount());
        System.out.println("misses        = " + s.missCount());
        System.out.printf("hitRate       = %.4f%n", s.hitRate());
        System.out.println("dbHits        = " + dao.dbHits());
        System.out.println("loadSuccess   = " + s.loadSuccessCount());
        System.out.println("avgLoadMs     = "
                + String.format("%.3f", s.totalLoadTime() / 1_000_000.0 / s.loadSuccessCount()));

        // Eviction: a 50-slot cache force-fed 200 products.
        Cache<Long, Product> small = Caffeine.newBuilder()
                .maximumSize(50)
                .recordStats()
                .build();
        for (long id = 1; id <= 200; id++) {
            small.put(id, dao.findById(id));
        }
        small.cleanUp(); // run pending maintenance so eviction is observable
        System.out.println("--- eviction ---");
        System.out.println("inserted      = 200");
        System.out.println("estimatedSize = " + small.estimatedSize());
        System.out.println("evictions     = " + small.stats().evictionCount());
    }
}

This was compiled against caffeine-3.1.8.jar with javac on Java 21 and actually run. The output below is the genuine console output from that run — not a mock:

requests      = 520
hits          = 500
misses        = 20
hitRate       = 0.9615
dbHits        = 20
loadSuccess   = 20
avgLoadMs     = 2.329
--- eviction ---
inserted      = 200
estimatedSize = 50
evictions     = 150

Read it like a production dashboard. Of 520 lookups, only 20 reached the DAO — the cache absorbed 500 database calls it would otherwise have made, a hit rate of 0.9615. loadSuccessCount (20) matches dbHits (20): every miss loaded exactly once, no duplicates. The eviction half proves the size bound is real: 200 products forced into a 50-slot cache, and after maintenance the cache holds exactly 50 with 150 evictions recorded. (The avgLoadMs figure is the measured mean of that run's 20 loads — it varies a little between runs; the counts are the stable facts.) Two details worth noticing: recordStats() is what makes stats() meaningful — without it every counter reads zero — and cleanUp() forces pending maintenance so the eviction counts are observable immediately instead of on the next read or write.

Decision rule: always build caches with recordStats() and expose the hit rate to your monitoring. A cache whose hit rate you cannot see is a cache whose value you cannot prove — and whose failure mode (silently becoming a second database call on every request) you will discover during the next flash sale.

Eviction: the cache is a budget, not a warehouse

An unbounded cache is a memory leak with good marketing. Caffeine gives you four ways to bound what it holds, and they answer different questions:

PolicyWhat it boundsWhen it fits
maximumSize(n)Entry count, evicted by W-TinyLFU (frequency + recency)Uniform values — product prices, session flags. The default choice.
maximumWeight(n) + weigherMemory, by a weight you define (e.g. byte length)Variable-size values — thumbnails, rendered pages, JSON blobs.
expireAfterWrite(d)Age since the value was writtenData with a known freshness contract — prices (10 min), inventory snapshots.
expireAfterAccess(d)Idle time since last read"Recently used" working sets — recently-viewed products. Rarely right for prices.
refreshAfterWrite(d)Staleness without blocking readersHot keys where a slow reload must never stall a request (see below).

The weight-based form, for when values vary wildly in size (a 2 KB price record and a 2 MB product image should not cost the same slot):

// Weight-based eviction: bound memory by bytes, not entry count.
Cache<Long, byte[]> thumbnails = Caffeine.newBuilder()
        .maximumWeight(64 * 1024 * 1024) // 64 MiB of thumbnails
        .weigher((Long id, byte[] bytes) -> bytes.length)
        .recordStats()
        .build();

And refreshAfterWrite deserves its own paragraph, because it solves a subtle problem: with plain expireAfterWrite, the first request after expiry pays the full reload cost — a latency spike on exactly the hottest key. refreshAfterWrite instead serves the stale value immediately and reloads asynchronously behind it. Readers never block on a refresh:

// refreshAfterWrite: stale reads never block; an async reload
// refreshes the entry behind the scenes.
LoadingCache<Long, Product> products = Caffeine.newBuilder()
        .maximumSize(10_000)
        .expireAfterWrite(Duration.ofMinutes(30))   // hard freshness ceiling
        .refreshAfterWrite(Duration.ofMinutes(5))   // soft refresh, non-blocking
        .recordStats()
        .build(dao::findById);

Note the pairing: refreshAfterWrite without an expireAfterWrite ceiling means a value whose reload keeps failing stays stale forever — the refresh only replaces the entry on success. The snippet above compiles against Caffeine 3.1.8 (verified); the async-refresh behavior is Caffeine's documented contract, not something this demo measured.

Decision rule: bound every cache by size or weight — never ship an unbounded one — and give time-sensitive data an expireAfterWrite that matches the business freshness contract (minutes for prices, hours for catalog copy), not a number that felt safe. Add refreshAfterWrite only for hot keys where one slow reload would spike p99 latency.

Going distributed: Redis, because one JVM is not the fleet

Caffeine is per-JVM. The moment checkout runs on eight instances, you have eight independent caches: eight cold starts, eight times the memory, and — the nasty one — eight slightly different views of the price at any moment. Instance A refreshed its price at 10:00, instance B at 10:04; a customer retrying a failed checkout can land on the other instance and see a different price. A distributed cache gives the whole fleet one shared, coherent view. That is what Redis is for here: not as a database, but as the fleet's shared price list.

Which client — Jedis or Lettuce? The two serious options are Jedis (redis.clients:jedis:6.1.0) and Lettuce (io.lettuce:lettuce-core). Lettuce is the more operationally complete client: non-blocking I/O, automatic reconnection, and Redis Cluster topology refresh. Jedis is synchronous and blocking — one thread, one connection, one command at a time. For this post I chose Jedis, deliberately: its blocking API maps one-to-one onto the cache-aside pattern you just learned, so every line below is about caching, not about reactive plumbing. If your service is already on WebFlux, or you need cluster failover handling, pick Lettuce then — the caching patterns transfer unchanged.

Serialization, decided up front: Redis stores bytes; your cache stores products. The three honest options are JSON (readable, language-agnostic, tolerates field additions), Java native serialization (never — fragile across deploys and a known deserialization-attack surface), and a binary format like Protobuf (fastest, but schema machinery for a price object is overkill). This demo uses JSON via Gson; in a Spring Boot service you would use the Jackson ObjectMapper already on your classpath instead of adding Gson.

The code below is the same cache-aside shape as the Caffeine demo, now against a shared Redis — plus explicit invalidation, which is how the Diwali price-drop incident gets prevented. It was compiled against jedis-6.1.0.jar (plus gson-2.11.0.jar and commons-pool2-2.12.0.jar, Jedis's pool dependency), and every Jedis method used was verified to exist in that jar version:

package com.javamakeuse.catalog;

import com.google.gson.Gson;
import java.time.Duration;
import java.util.UUID;
import redis.clients.jedis.Jedis;
import redis.clients.jedis.JedisPool;
import redis.clients.jedis.JedisPoolConfig;
import redis.clients.jedis.params.SetParams;

/**
 * Cache-aside against a shared Redis, using Jedis. Every checkout
 * instance in the fleet reads the same prices; one TTL bounds staleness.
 *
 * COMPILE-ONLY: verified against jedis-6.1.0.jar but never executed —
 * there is no Redis server in this environment.
 */
public class RedisProductCache implements AutoCloseable {

    private static final Duration PRICE_TTL = Duration.ofMinutes(5);

    private final JedisPool pool;
    private final FakeProductDao dao;
    private final Gson gson = new Gson();

    public RedisProductCache(String host, int port, FakeProductDao dao) {
        JedisPoolConfig config = new JedisPoolConfig();
        config.setMaxTotal(32);
        config.setMaxIdle(16);
        this.pool = new JedisPool(config, host, port);
        this.dao = dao;
    }

    private static String key(long productId) {
        return "catalog:product:" + productId;
    }

    /** Cache-aside: Redis first, database on miss, write-back with TTL. */
    public Product getProduct(long productId) {
        String k = key(productId);
        try (Jedis jedis = pool.getResource()) {
            String json = jedis.get(k);
            if (json != null) {
                return gson.fromJson(json, Product.class);
            }
        }
        Product fresh = dao.findById(productId);
        try (Jedis jedis = pool.getResource()) {
            jedis.set(k, gson.toJson(fresh),
                    SetParams.setParams().ex(PRICE_TTL.toSeconds()));
        }
        return fresh;
    }

    /** Explicit invalidation: the price service calls this the moment
     *  a price changes, instead of waiting out the TTL. */
    public void onPriceChanged(long productId) {
        try (Jedis jedis = pool.getResource()) {
            jedis.del(key(productId));
        }
    }

    @Override
    public void close() {
        pool.close();
    }

Three things to steal from this listing. First, the connection discipline: a JedisPool created once, and every use borrows a connection in try-with-resources — a leaked Jedis connection is a pool slot gone until restart, and an exhausted pool looks exactly like a database outage. Second, the key namespace (catalog:product:<id>): prefix every key with its domain, or the day someone else's session keys collide with your product keys you will debug ghosts. Third, onPriceChanged: the price-update path deletes the key the moment the database row changes. That single del is the entire difference between the two Diwali incidents.

Honesty note: I could not run this listing — no Redis server exists in this environment, and I will not fabricate console output. Against a local redis-server, the behavior to verify is: first getProduct(7) misses, loads from the DAO, and stores JSON with a 300-second TTL; the second call returns without touching the DAO; after onPriceChanged(7), redis-cli GET catalog:product:7 returns nil and the next read reloads. Run those three checks before trusting any Redis cache in production.

Decision rule: use Caffeine when one JVM's hot set fits in memory and slight cross-instance staleness is tolerable; add Redis when the fleet must agree on the value, when the hot set exceeds one heap, or when a restart must not mean a cold cache. Many production checkouts run both: Caffeine as L1 (microseconds, no network) in front of Redis as L2 (shared, milliseconds).

The thundering herd: when the cache is cold at the worst moment

There is a failure mode the cache-aside diagram hides. Picture a deploy at 11:55 AM: all eight checkout instances restart, every Caffeine cache is empty, and the flash-sale traffic is still coming. The first wave of requests all miss at the same instant — and every one of them runs the loader against the database simultaneously. The cache you built to protect the database becomes the mechanism that synchronizes the attack on it. This is the cache stampede (or thundering herd), and the naive check-then-act code is the loaded gun:

// BROKEN under concurrency: the gap between the check and the put
// is where the herd runs through.
String price = cache.getIfPresent(key);
if (price == null) {
    price = backend.loadPrice(key); // 64 threads all do this at once
    cache.put(key, price);
}

The demo below proves it. Sixty-four threads, released at the same instant by a latch, all ask for the same cold product; the backend takes 200 ms per load and counts invocations. First the naive pattern, then Caffeine's get(key, loader) — whose mapping function Caffeine guarantees to run at most once per key, so the other 63 threads simply wait for the one in-flight load (the singleflight pattern, built in):

package com.javamakeuse.catalog;

import com.github.benmanes.caffeine.cache.Cache;
import com.github.benmanes.caffeine.cache.Caffeine;
import java.time.Duration;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.CountDownLatch;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.atomic.AtomicInteger;

/**
 * The thundering herd, demonstrated. 64 threads ask for the same cold
 * product at the same instant. The naive check-then-load stampedes the
 * "database"; Caffeine's per-key atomic load lets exactly one thread in.
 */
public class StampedeDemo {

    private static final int THREADS = 64;
    private static final long HOT_KEY = 7L;

    // A slow backend call: 200 ms, and it counts invocations.
    static class SlowBackend {
        final AtomicInteger loads = new AtomicInteger();

        String loadPrice(long productId) {
            loads.incrementAndGet();
            try {
                Thread.sleep(200);
            } catch (InterruptedException e) {
                Thread.currentThread().interrupt();
            }
            return "price-of-" + productId;
        }
    }

    public static void main(String[] args) throws Exception {
        ExecutorService pool = Executors.newFixedThreadPool(THREADS);
        try {
            naiveStampede(pool);
            caffeineSingleflight(pool);
        } finally {
            pool.shutdown();
        }
    }

    /** Broken pattern: getIfPresent + load + put. The gap between the
     *  check and the put is where the herd runs through. */
    private static void naiveStampede(ExecutorService pool) throws Exception {
        Map<Long, String> cache = new ConcurrentHashMap<>();
        SlowBackend backend = new SlowBackend();
        CountDownLatch start = new CountDownLatch(1);
        CountDownLatch done = new CountDownLatch(THREADS);

        for (int i = 0; i < THREADS; i++) {
            pool.submit(() -> {
                try {
                    start.await();
                    String price = cache.get(HOT_KEY);
                    if (price == null) {
                        price = backend.loadPrice(HOT_KEY);
                        cache.put(HOT_KEY, price);
                    }
                } catch (InterruptedException e) {
                    Thread.currentThread().interrupt();
                } finally {
                    done.countDown();
                }
                return null;
            });
        }
        long t0 = System.nanoTime();
        start.countDown();
        done.await();
        long ms = (System.nanoTime() - t0) / 1_000_000;
        System.out.println("[naive]    backend loads = " + backend.loads.get()
                + " for " + THREADS + " concurrent requests (wall ~" + ms + " ms)");
    }

    /** Caffeine: the mapping function runs at most once per key, no
     *  matter how many threads arrive together. */
    private static void caffeineSingleflight(ExecutorService pool) throws Exception {
        SlowBackend backend = new SlowBackend();
        Cache<Long, String> cache = Caffeine.newBuilder()
                .maximumSize(1_000)
                .expireAfterWrite(Duration.ofMinutes(10))
                .build();
        CountDownLatch start = new CountDownLatch(1);
        CountDownLatch done = new CountDownLatch(THREADS);

        for (int i = 0; i < THREADS; i++) {
            pool.submit(() -> {
                try {
                    start.await();
                    cache.get(HOT_KEY, k -> backend.loadPrice(k));
                } catch (InterruptedException e) {
                    Thread.currentThread().interrupt();
                } finally {
                    done.countDown();
                }
                return null;
            });
        }
        long t0 = System.nanoTime();
        start.countDown();
        done.await();
        long ms = (System.nanoTime() - t0) / 1_000_000;
        System.out.println("[caffeine] backend loads = " + backend.loads.get()
                + " for " + THREADS + " concurrent requests (wall ~" + ms + " ms)");
    }
}

Genuine output from the run (Java 21, compiled against caffeine-3.1.8.jar):

[naive]    backend loads = 64 for 64 concurrent requests (wall ~256 ms)
[caffeine] backend loads = 1 for 64 concurrent requests (wall ~208 ms)

The naive version fired 64 backend loads for 64 requests — the "cache" absorbed nothing. Caffeine's version fired exactly 1. (Wall-clock times are from this machine and will vary; the load counts are the point, and they are structural: Caffeine computes the mapping function atomically per key.) The lesson generalizes beyond Caffeine: never write check-then-act around a cache under concurrency — use the cache's atomic load operation (get(key, loader), or ConcurrentHashMap.computeIfAbsent for plain maps).

Across a fleet, one JVM's singleflight is not enough — eight instances can still stampede together. The distributed fix is a short-lived lock in Redis: the first instance to miss takes the lock with SET key token NX EX 30 (set only if absent, auto-expiring so a crashed holder cannot wedge the key forever); the losers poll the cache briefly instead of hitting the database. The lock release must be a Lua compare-and-del — deleting unconditionally could remove a new lock taken by another instance after yours expired. This is the second half of the compiled RedisProductCache:

    /**
     * Distributed singleflight for a fleet: the first instance to miss
     * takes a short-lived lock (SET key token NX EX); the losers wait
     * and re-read instead of stampeding the database.
     */
    public Product getProductSingleflight(long productId) throws InterruptedException {
        Product hit = readThroughOnce(productId);
        if (hit != null) {
            return hit;
        }
        String lockKey = key(productId) + ":lock";
        String token = UUID.randomUUID().toString();
        try (Jedis jedis = pool.getResource()) {
            String acquired = jedis.set(lockKey, token,
                    SetParams.setParams().nx().ex(30));
            if (!"OK".equals(acquired)) {
                // Someone else is loading it: poll the cache briefly.
                for (int i = 0; i < 40; i++) {
                    Thread.sleep(50);
                    Product retry = readThroughOnce(productId);
                    if (retry != null) {
                        return retry;
                    }
                }
                return readThroughOnce(productId); // last resort
            }
            try {
                Product fresh = dao.findById(productId);
                try (Jedis j2 = pool.getResource()) {
                    j2.set(key(productId), gson.toJson(fresh),
                            SetParams.setParams().ex(PRICE_TTL.toSeconds()));
                }
                return fresh;
            } finally {
                // Release only if we still own the lock (compare-and-del).
                jedis.eval("if redis.call('get', KEYS[1]) == ARGV[1] "
                                + "then return redis.call('del', KEYS[1]) else return 0 end",
                        java.util.List.of(lockKey), java.util.List.of(token));
            }
        }
    }

    private Product readThroughOnce(long productId) {
        try (Jedis jedis = pool.getResource()) {
            String json = jedis.get(key(productId));
            return json == null ? null : gson.fromJson(json, Product.class);
        }
    }
}

Decision rule: in-process, always load through the atomic get(key, loader) — never check-then-act. Distributed, guard cold-key loads with a SET NX EX lock plus Lua compare-and-del release, and keep the lock TTL short (tens of seconds): it only needs to cover one load, and a shorter TTL bounds the damage if a holder dies.

TTL vs explicit invalidation: pick your staleness contract

Every cached value is a promise about staleness, and there are exactly two ways to keep it:

TTL (time-to-live)Explicit invalidation
How freshness happensEntries die of old age; the next read reloadsA writer deletes/updates the key the moment data changes
Worst-case stalenessBounded by the TTL — and you chose itZero, if every writer cooperates
Failure modeServes stale data for up to TTL after a changeOne writer that forgets to invalidate = stale forever
Stampede riskReal: hot keys expire together and reload togetherLow: invalidations are driven by writes, which are rarer
EffortOne line: ex(...)Every write path must know about the cache

In practice you combine them, which is what this post's code does: a 5-minute TTL as the backstop (staleness is never worse than 5 minutes, even if a bug skips invalidation) plus onPriceChanged for immediacy (price changes propagate in milliseconds on the paths you control). Belt and suspenders, and each covers the other's failure mode.

Decision rule: TTL alone for data where bounded staleness is harmless (catalog copy, category trees — pick the TTL from the business contract, not from habit). Invalidation on top the moment staleness costs money (prices, inventory counts, feature flags). If you cannot enumerate every writer of the data, you cannot do explicit invalidation — admit it and choose the TTL accordingly.

When NOT to cache

Caching is not a performance strategy; it is a correctness-risky performance tactic. Do not cache when:

  • The data is per-user with no reuse. A cache hit rate near zero is just a slower, more complicated database call with extra failure modes. Check stats().hitRate() — if it sits below ~0.5 for your workload, delete the cache, don't tune it.
  • Staleness costs money and you have no invalidation hook. Prices written by a third-party feed you can't hook into, balances, entitlements — if you can't invalidate on write, a cache is a bug waiting for a refund ticket.
  • You're caching to hide a broken query. If the price lookup is slow because of an N+1 or a missing index, caching makes the symptom intermittent instead of fixing it — and the first cold start reminds you. Fix the query first (see "JPA Performance: LAZY/EAGER, N+1, Fetch Joins & Locking"), then cache the fast query.
  • The dataset fits in the database's buffer pool anyway. A 10,000-row product table lives in Postgres's shared buffers; your "optimization" adds network hops and serialization to reach data the DB serves from RAM.
  • You would cache errors. Caching a failed lookup ("product not found") without a deliberate, short, separate TTL turns a transient outage into a persistent one. Negative caching is a conscious choice with its own TTL — never the default path.

Decision rule: cache read-heavy, slowly-changing, widely-shared data with a measurable hit rate — and write the invalidation story before the caching story. If you can't describe, in one sentence, what makes an entry stale and who removes it, you aren't ready to cache it.

What's next

The checkout now reads prices at memory speed without melting the database or serving yesterday's prices. But not every job belongs in the request path: nightly price-list imports, settlement reports, and email digests are batch work, and they deserve batch machinery. The next post in this track is "Batch Jobs with Spring Batch": chunk-oriented steps, restartability after failure, and why you should never loop over a million rows in a controller.

Field check before you move on: take one read-heavy endpoint you own and wrap it in a Caffeine cache with recordStats(), a maximumSize, and an expireAfterWrite you can defend to the data owner. Run your normal load against it for a day and read stats(): hit rate, eviction count, average load time. Then simulate the herd — restart the service under load and watch whether the database spikes. If it does, convert the load path to the atomic get(key, loader) form (or the Redis lock form for a fleet) and re-run. The before/after stats are the whole lesson, measured on your own system.

Continue: Java Learning Roadmap 2026

Comments

Popular posts from this blog

JSP Servlet Interview Questions For Freshers Series 1

Java Banking Finance Services and Insurance (BFSI) domain interview questions

Java program to check even or odd number