Virtual Threads: One Per Task

The flash sale started at midnight. By 12:40 a.m. the order-intake service was gasping: its fixed pool of 200 platform threads was saturated. Every request spent ~150ms blocked — calling the inventory service, then the pricing service, then the fraud check — so 200 threads × (1000ms / 150ms) capped the service at roughly 1,300 requests per second, and the sale was pushing twice that. Requests queued, timeouts cascaded, and the on-call engineer's "fix" was the obvious one: raise the pool to 8,000 threads. The JVM died within minutes with OutOfMemoryError: unable to create native thread.

The team had treated thread scarcity as the problem. It wasn't. Each platform thread is an OS thread with roughly a megabyte of reserved stack, and the operating system simply refused to create the eight-thousandth-and-first. The real problem was the programming model: blocking code needed a thread per request, and threads were too expensive to have one per request. Virtual threads — final in JDK 21 — break that tradeoff. They make blocking cheap, so you can go back to writing straightforward, one-thread-per-task code at a scale platform threads could never reach.

This post is hands-on. Every number below comes from programs I compiled and ran on JDK 21.0.3 (a 2-CPU box, which matters for a couple of the timings — I'll say where). You'll see 100,000 virtual threads do useful work where 8,023 platform threads killed the JVM, and you'll learn the one principle that keeps virtual threads from becoming a new way to melt your database.

1. Why platform threads are scarce

A platform thread is a one-to-one wrapper around an operating-system thread, and the OS charges rent: each thread reserves stack memory up front. On a 64-bit JVM the default reservation is 1MB per thread (-Xss1m). That reservation is virtual address space, not physical RAM, but the OS still enforces limits on how many native threads a process may create.

The arithmetic is brutal. 100,000 platform threads × 1MB = ~100GB of reserved address space, before a single byte of your application data. Let's watch a JVM try. This program creates platform threads in a loop until the OS says no:

import java.util.ArrayList;

public class PlatformLimit {
    public static void main(String[] args) {
        ArrayList<Thread> keep = new ArrayList<>();
        int created = 0;
        try {
            while (true) {
                Thread t = Thread.ofPlatform().unstarted(() -> {
                    try { Thread.sleep(60_000); } catch (InterruptedException e) {}
                });
                t.setDaemon(true);
                t.start();
                keep.add(t);
                created++;
            }
        } catch (Throwable e) {
            System.out.println("FAILED after creating " + created
                + " platform threads: " + e);
        }
    }
}
[18.911s][warning][os,thread] Failed to start thread "Unknown thread" - pthread_create failed (EAGAIN) for attributes: stacksize: 1024k, guardsize: 0k, detached.
FAILED after creating 8023 platform threads: java.lang.OutOfMemoryError: unable to create native thread: possibly out of memory or process/resource limits reached

Read that carefully: 8,023 platform threads and the JVM is done — on a machine with gigabytes of free heap. The heap was never the problem; native thread creation was. (The warning line confirms the 1MB-per-thread stack reservation: stacksize: 1024k.) This is exactly the wall the flash-sale team hit. And it's why, for twenty years, Java developers learned elaborate workarounds — fixed thread pools sized by guesswork, async callbacks, reactive pipelines — all to avoid needing "too many" threads.

Virtual threads change the price of a thread so dramatically that the workarounds stop being necessary. A virtual thread is a JVM-managed thread: cheap to create (microseconds), tiny when idle, and — the key insight — it doesn't hold a platform thread while it's blocked. Keep that sentence; section 3 shows the mechanism.

2. Virtual threads: one per task

Creating one is a single expression, and the style will look familiar — that's the point:

public class WhoAmI {
    public static void main(String[] args) throws Exception {
        Thread vt = Thread.ofVirtual().name("order-worker").unstarted(() -> {
            System.out.println("inside : " + Thread.currentThread());
            System.out.println("virtual? " + Thread.currentThread().isVirtual());
        });
        vt.start();
        vt.join();
        System.out.println("carrier pool parallelism (this box): "
            + Runtime.getRuntime().availableProcessors());
    }
}
inside : VirtualThread[#18,order-worker]/runnable@ForkJoinPool-1-worker-1
virtual? true
carrier pool parallelism (this box): 2

Three things to notice. First, the thread has a name you chose (order-worker) — name your virtual threads; thread dumps with a million Thread-48291 entries are useless. Second, isVirtual() returns true — old APIs like Thread.currentThread() keep working, so existing libraries mostly behave. Third, the funny suffix: @ForkJoinPool-1-worker-1. That is the carrier — the platform thread currently executing this virtual thread. The JVM maintains a small pool of carriers (by default, one per CPU; this box has 2) and mounts virtual threads onto them. When a virtual thread blocks, it unmounts, freeing the carrier for another virtual thread. Section 3 is all about that handoff.

For the common "one thread per request" shape, reach for the thread-per-task executor instead of managing threads by hand:

import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;

public class VtExecutor {
    public static void main(String[] args) throws Exception {
        try (ExecutorService exec = Executors.newThreadPerTaskExecutor(
                Thread.ofVirtual().name("req-", 0).factory())) {
            for (int i = 1; i <= 5; i++) {
                final int id = i;
                exec.submit(() -> {
                    System.out.println("request " + id + " on "
                        + Thread.currentThread().getName());
                    return null;
                });
            }
        } // close() waits for every submitted task to finish
        System.out.println("all done; main was " + Thread.currentThread().getName());
    }
}
request 2 on req-1
request 1 on req-0
request 3 on req-2
request 4 on req-3
request 5 on req-4
all done; main was main

Two conventions worth adopting now. Thread.ofVirtual().name("req-", 0).factory() gives every thread a sequenced name (req-0, req-1, …) — free observability. And the try-with-resources block isn't just tidy: closing the executor waits for all submitted tasks, so main can't exit early and silently abandon work. (The request lines print in a different order on every run — that's genuine concurrency, not a bug in the demo. If your tests assert on output order of concurrent tasks, the tests are wrong.)

One more habit to unlearn: do not pool virtual threads. Pools exist to amortize the cost of creating platform threads. Virtual threads are cheap enough to create per task — pooling them just reintroduces the queueing, sizing, and rejection policies you were trying to escape. newThreadPerTaskExecutor with virtual threads is the replacement for the tuned ThreadPoolExecutor you used to agonize over.

One default that bites people exactly once: virtual threads are daemon threads, always — verified on this JVM:

public class VtDaemon {
    public static void main(String[] args) {
        Thread vt = Thread.ofVirtual().unstarted(() -> {});
        Thread pt = Thread.ofPlatform().unstarted(() -> {});
        System.out.println("virtual thread daemon by default? " + vt.isDaemon());
        System.out.println("platform thread daemon by default? " + pt.isDaemon());
    }
}
virtual thread daemon by default? true
platform thread daemon by default? false

A daemon thread can't keep the JVM alive: if main finishes while virtual threads are still running, the JVM exits and takes them with it — silently. (You also can't change it: setDaemon(false) on a virtual thread throws.) That's why the try-with-resources executor in the earlier example matters so much: it's not tidiness, it's the thing standing between your tasks and an early JVM exit. In long-lived servers this never comes up — something non-daemon is always running — but in CLI tools, batch jobs, and tests, always join or use a scoped executor.

3. Why blocking is cheap: mount and unmount

Here's the mechanism that makes everything else possible. When a platform thread blocks — Thread.sleep, a socket read, waiting on a lock — its OS thread sits parked and useless. Nothing else can use it. That's why 200 blocked platform threads meant 200 dead CPUs' worth of capacity in the flash-sale incident.

When a virtual thread blocks, the JVM unmounts it from its carrier: the virtual thread's stack is moved to the heap, the carrier platform thread goes back to the scheduler's pool, and another virtual thread mounts in its place. When the blocking operation completes, the virtual thread is rescheduled onto any free carrier. The carrier threads — one per CPU by default — stay busy doing real work while thousands of virtual threads wait. Blocking went from "holding an expensive resource hostage" to "parking cheap state on the heap."

100,000 virtual threads cheap: stack lives on the heap VT-1: RUNNING (mounted) VT-2: sleeping (unmounted) VT-3: socket read (unmounted) VT-4 … VT-100000 waiting blocked = parked on heap, carrier freed 2 carrier threads ForkJoinPool, 1 per CPU carrier-1 now running VT-1 carrier-2 now running VT-4 never idle while work waits

Let's measure it. One hundred thousand virtual threads, each sleeping 100ms — the purest "blocked task" workload there is. Platform threads died at 8,023; watch what virtual threads do:

public class VtScale2 {
    static long run(int n, long sleepMs) throws Exception {
        Thread[] threads = new Thread[n];
        long start = System.nanoTime();
        for (int i = 0; i < n; i++) {
            final long s = sleepMs;
            threads[i] = Thread.ofVirtual().unstarted(() -> {
                try { Thread.sleep(s); } catch (InterruptedException e) {
                    Thread.currentThread().interrupt();
                }
            });
            threads[i].start();
        }
        for (Thread t : threads) t.join();
        return (System.nanoTime() - start) / 1_000_000;
    }
    public static void main(String[] args) throws Exception {
        System.out.println("n=1000   sleep=0   wall=" + run(1000, 0) + "ms");
        System.out.println("n=10000  sleep=0   wall=" + run(10000, 0) + "ms");
        System.out.println("n=100000 sleep=0   wall=" + run(100000, 0) + "ms");
        System.out.println("n=10000  sleep=100 wall=" + run(10000, 100) + "ms");
        System.out.println("n=100000 sleep=100 wall=" + run(100000, 100) + "ms");
    }
}
n=1000   sleep=0   wall=72ms
n=10000  sleep=0   wall=148ms
n=100000 sleep=0   wall=1412ms
n=10000  sleep=100 wall=870ms
n=100000 sleep=100 wall=4940ms

100,000 virtual threads ran to completion on a 2-CPU box — the same box where 8,023 platform threads killed the JVM. But be honest about what the numbers say: the wall time is not "just over 100ms." Look at the sleep=0 column: launching 100,000 threads from a single loop costs ~1.4 seconds here (~14µs per start — thread creation is cheap, not free). The main thread starts them one at a time, so the last thread starts ~1.4s after the first and wakes ~100ms after its own start. Total wall time ≈ launch spread + sleep + scheduling on just 2 carriers.

That's the correct mental model: virtual threads remove the ceiling on how many blocked tasks you can have in flight, and per-thread startup is microseconds — but launching a million of them in a hot loop still takes real time, and a 2-CPU scheduler still schedules. On a bigger machine with more carriers these numbers shrink; the shape of the result doesn't change. The flash-sale service's 200-thread cap wasn't a law of nature. It was a price tag.

4. The signature principle: threads are not connections

Here is the sentence this whole post exists to teach:

"virtual threads remove thread scarcity, not resource scarcity" — 1M virtual threads ≠ 1M DB connections; you must still bound downstream resources (e.g., Semaphore around a connection pool).

Virtual threads make it trivially easy to have 100,000 tasks in flight. Your database cannot serve 100,000 concurrent queries. Your downstream HTTP service cannot take 100,000 concurrent calls. The thread was never the scarce resource — the connection was, and virtual threads don't create connections. The failure mode just moved: instead of "can't create threads," you get "pool exhausted, timeouts, cascading failure" — now at a scale that can take down the database for everyone, not just your service.

The fix is old and boring, which is why it works: bound the scarce resource explicitly with a Semaphore, sized to the downstream capacity (usually your connection pool size). This demo simulates it — a fake "DB pool" of 8 connections, 1,000 incoming requests as virtual threads, each "query" taking 20ms:

import java.util.concurrent.Semaphore;
import java.util.concurrent.atomic.AtomicInteger;

public class DbPool {
    static final int POOL = 8;      // the fake DB connection pool has 8 connections
    static final int TASKS = 1000;  // 1000 incoming requests
    static final Semaphore pool = new Semaphore(POOL);
    static final AtomicInteger inUse = new AtomicInteger();
    static final AtomicInteger maxInUse = new AtomicInteger();

    public static void main(String[] args) throws Exception {
        Thread[] vt = new Thread[TASKS];
        long start = System.nanoTime();
        for (int i = 0; i < TASKS; i++) {
            vt[i] = Thread.ofVirtual().unstarted(() -> {
                try {
                    pool.acquire();                       // wait for a connection
                    int cur = inUse.incrementAndGet();
                    maxInUse.accumulateAndGet(cur, Math::max);
                    Thread.sleep(20);                     // the "query"
                    inUse.decrementAndGet();
                    pool.release();
                } catch (InterruptedException e) {
                    Thread.currentThread().interrupt();
                }
            });
            vt[i].start();
        }
        for (Thread t : vt) t.join();
        long ms = (System.nanoTime() - start) / 1_000_000;
        System.out.println("BOUNDED   tasks=" + TASKS + " pool=" + POOL
            + " wall=" + ms + "ms maxConcurrentQueries=" + maxInUse.get());
    }
}
BOUNDED   tasks=1000 pool=8 wall=3109ms maxConcurrentQueries=8

Read the evidence: maxConcurrentQueries=8 — across 1,000 tasks, never 9 concurrent "queries." The gate held perfectly. And the wall time tells the same story from the other side: 1,000 tasks ÷ 8 permits = 125 waves × 20ms = 2,500ms of pure query time, and we measured 3,109ms (the rest is thread-launch spread and scheduling on 2 carriers, as section 3 taught us). Throughput was gated by the 8 permits, exactly as designed.

Now the same 1,000 tasks with no bound at all:

public class NoPool {
    public static void main(String[] args) throws Exception {
        int tasks = 1000;
        Thread[] vt = new Thread[tasks];
        long start = System.nanoTime();
        for (int i = 0; i < tasks; i++) {
            vt[i] = Thread.ofVirtual().unstarted(() -> {
                try { Thread.sleep(20); } catch (InterruptedException e) {
                    Thread.currentThread().interrupt();
                }
            });
            vt[i].start();
        }
        for (Thread t : vt) t.join();
        long ms = (System.nanoTime() - start) / 1_000_000;
        System.out.println("UNBOUNDED tasks=" + tasks + " wall=" + ms + "ms");
    }
}
UNBOUNDED tasks=1000 wall=277ms

The unbounded version is eleven times faster — and that's precisely the trap. It's fast here because the "query" is fake: nothing real is being consumed. Point those 1,000 unbound virtual threads at a real database and you don't get 277ms; you get 1,000 simultaneous connections against a pool sized for 50, connection timeouts, retry storms, and an incident that pages the database team. (I describe this rather than demo it: actually opening 1,000 sockets to prove a point would be vandalism, not engineering.)

1,000 virtual threads requests arrive as fast as they like VT-1 … VT-1000 waiting is cheap unmounted while waiting for a permit Semaphore(8) 8 permits = 8 connections the gate Database max 8 concurrent queries — never 9 pool stays healthy latency stays flat measured: 3109ms wall for 1000 queries

The pattern to take to production: size the semaphore to the downstream resource (your HikariCP maximumPoolSize, the downstream service's documented concurrency limit), acquire() before touching the resource, release() in a finally. A virtual thread blocked on acquire() unmounts — it costs you nothing while it waits. The threads queue; the database doesn't. That is what "remove thread scarcity, not resource scarcity" means in code.

5. Structured concurrency, briefly

One-per-task threading creates a bookkeeping problem: if a method fans out to three virtual threads and one fails, who cancels the other two? Who waits for stragglers before returning? Structured concurrency answers with StructuredTaskScope: subtasks are children of a scope, and the scope's try-with-resources block doesn't exit until every child is done. Failure handling becomes declarative instead of scattered.

One honest flag first: in JDK 21 StructuredTaskScope is a preview API — the demo below compiles and runs with --enable-preview (it was finalized in a later JDK). The idea is what matters here; the API shape is stable enough to learn.

import java.util.concurrent.StructuredTaskScope;

public class ScopeDemo2 {
    static String fetchUser() throws InterruptedException {
        Thread.sleep(150); return "user:ada";
    }
    static String fetchOrders() throws InterruptedException {
        Thread.sleep(250); return "orders:3";
    }

    static long fetchBoth() throws Exception {
        long start = System.nanoTime();
        try (var scope = new StructuredTaskScope.ShutdownOnFailure()) {
            var user = scope.fork(ScopeDemo2::fetchUser);
            var orders = scope.fork(ScopeDemo2::fetchOrders);
            scope.join();            // wait for both
            scope.throwIfFailed();   // if either threw, rethrow here
            System.out.println(user.get() + " | " + orders.get());
        }
        return (System.nanoTime() - start) / 1_000_000;
    }
    public static void main(String[] args) throws Exception {
        fetchBoth(); // warmup
        System.out.println("warmed-up elapsed=" + fetchBoth()
            + "ms (sequential would be ~400ms)");
    }
}
user:ada | orders:3
user:ada | orders:3
warmed-up elapsed=264ms (sequential would be ~400ms)

The two fetches overlapped — elapsed is near the slower of the two (250ms) rather than their sum (400ms). ShutdownOnFailure adds the failure half of the contract: if one subtask throws, the scope shuts down the others and throwIfFailed() rethrows the first failure — no orphaned threads, no half-fetched results leaking past the method. That's the whole pitch, and deliberately so: structured concurrency is a small idea. Use it wherever you'd otherwise fork tasks and hand-roll the join-and-cancel logic.

6. Caveats: pinning and ThreadLocal

Two sharp edges, stated plainly.

Carrier pinning. A virtual thread is pinned when the JVM cannot unmount it from its carrier — the carrier sits stuck for the whole block, and your "2 carriers" silently become 1, then 0. On JDK 21 through 23, the classic cause is blocking inside a synchronized block or method: the monitor is tied to the carrier thread, so a virtual thread that parks while holding it pins the carrier. (This changed in JDK 24 via JEP 491 — synchronized no longer pins there — but you're running 21, so the caveat applies to you.) Native/JNI calls can still pin on every version.

I'm not showing a pinning demo, and here's why: pinning is timing-sensitive — it shows up as degraded throughput under contention, not as a clean pass/fail you can print. A demo that "works on my machine" would teach you nothing reliable. The documented behavior is enough to act on: prefer ReentrantLock over synchronized when the critical section can block. ReentrantLock is implemented in Java and never pins:

import java.util.concurrent.locks.ReentrantLock;

public class VtLock {
    static final ReentrantLock lock = new ReentrantLock();

    public static void main(String[] args) throws Exception {
        Thread[] vt = new Thread[200];
        for (int i = 0; i < vt.length; i++) {
            vt[i] = Thread.ofVirtual().unstarted(() -> {
                lock.lock();               // ReentrantLock: carrier is NOT pinned
                try {
                    Thread.sleep(10);      // blocking call inside the lock
                } catch (InterruptedException e) {
                    Thread.currentThread().interrupt();
                } finally {
                    lock.unlock();
                }
            });
            vt[i].start();
        }
        for (Thread t : vt) t.join();
        System.out.println("200 virtual threads serialized through ReentrantLock, all finished");
    }
}
200 virtual threads serialized through ReentrantLock, all finished

Two hundred virtual threads serialized through one lock on 2 carriers, no drama — each blocked thread unmounted and its carrier went on to run the next waiter. If you suspect pinning in production, the JDK ships a flight recorder event for exactly this: jdk.VirtualThreadPinned. One JFR recording answers "is pinning my problem?" in minutes.

ThreadLocal at scale. ThreadLocal works fine in virtual threads — but do the multiplication before you lean on it. A ThreadLocal<byte[]> holding a 1MB buffer is harmless with 200 platform threads (200MB, already questionable) and catastrophic with a million virtual threads (a terabyte that doesn't exist). ThreadLocals also defeat the "cheap thread" economics: values linger for the thread's lifetime, and with thread-per-task, that lifetime is one request — which is either exactly what you want or a subtle leak, depending on whether anything ever calls remove(). Prefer passing context as method arguments; for request-scoped values shared across many methods, look at ScopedValue (also preview in 21) as the structured alternative. The rule of thumb: nothing in a virtual thread should scale with the thread count unless you've done the arithmetic.

And the meta-caveat that covers both: virtual threads are for blocking I/O-bound work — request handlers, fan-out calls, polling loops. CPU-bound work (number crunching, compression) gains nothing from unmounting; it still needs one carrier per core, and oversubscribing carriers just adds scheduling overhead. Match the tool to the workload.

7. Adopting virtual threads without a 2 a.m. page

Back to the flash-sale team. Their fix wasn't "more threads" — it was changing what a thread costs. With virtual threads, 5,000 concurrent blocked requests become 5,000 virtual threads sharing a handful of carriers: the OutOfMemoryError: unable to create native thread failure mode disappears entirely. But section 4's warning still applies — the database behind those requests didn't get any bigger. The team's real checklist looked like this, and it's the one I'd hand you:

1. Switch at the edges first. The lowest-risk adoption is the entry point: replace the fixed request-handling pool with Executors.newThreadPerTaskExecutor(Thread.ofVirtual().factory()), or in Spring Boot 3.2+, the single property spring.threads.virtual.enabled=true. One change, and every request handler becomes a cheap virtual thread. Measure before and after — p99 latency and thread count are the two numbers that tell you it worked.

2. Audit synchronized on hot paths. On JDK 21–23, blocking inside synchronized pins the carrier (section 6). Grep for synchronized in request-handling code; anywhere the critical section can block on I/O, switch to ReentrantLock. Library code you don't control gets the same treatment — check whether your HTTP client and connection pool are virtual-thread-friendly before you blame the scheduler.

3. Bound every downstream resource. This is the signature principle as an action item: for each database, queue, and downstream service, put a Semaphore (or a bounded pool) between your now-unbounded threads and the finite resource, sized to its capacity — not your thread count. The semaphore from section 4 isn't demo code; it's production code.

4. Name your threads. Thread.ofVirtual().name("order-handler-", 0) costs nothing and turns a thread dump with 50,000 anonymous entries into a readable one. Future-you, debugging at 2 a.m., will be grateful.

5. Audit ThreadLocal. Grep for it, and for each one do the multiplication: bytes held × peak virtual thread count. Anything that scales with thread count gets refactored into method arguments or a scoped value.

6. Load-test the downstream, not just your service. Virtual threads move the bottleneck — that's the good news and the warning. Before rollout, saturate a staging database with the new concurrency level and watch its metrics: active connections, lock waits, buffer-pool pressure. If the database was sized for 200 concurrent clients and you're about to send it 5,000, the migration plan includes a database conversation, not just a JVM flag.

7. Leave the carrier pool alone. The default scheduler sizes carriers to your CPU count, which is correct for I/O-bound work: carriers should equal the parallelism your hardware actually has. The system property jdk.virtualThreadScheduler.parallelism exists, but reaching for it is a smell — if carriers are saturated, you usually have CPU-bound work mislabeled as I/O-bound (or pinning; check JFR's jdk.VirtualThreadPinned event first).

Notice what isn't on the list: rewriting blocking code into async code. That's the whole point. The straightforward, blocking, one-thread-per-task style your team already knows how to write, review, and debug is now the scalable style — as long as the scarce resources behind it are explicitly bounded.

What's next

Virtual threads solved the "one thread per task" problem by making threads cheap — but they didn't change how you compose asynchronous work. Sometimes you don't want to block a thread at all: you want "fetch the user, then fetch their orders, then combine, with a timeout on the whole thing" expressed as a pipeline rather than nested blocking calls. That's CompletableFuture: Async Pipelines Without Callback Hell, the next post — where you'll learn to chain, combine, and time-box async stages without drowning in callbacks, and when to pick a future over a virtual thread.

Field check before you move on: take the DbPool demo from section 4 and change two numbers: raise TASKS to 10,000 and lower POOL to 4. Predict the wall time first — 10,000 ÷ 4 = 2,500 waves × 20ms = 50,000ms — then run it and compare. Next, add a third AtomicInteger that counts how many threads are waiting on pool.acquire() at the peak (increment before acquire, decrement after). You should see the waiting count climb into the thousands while maxInUse stays pinned at 4. That picture — thousands waiting cheaply, 4 working — is the entire post in one run. If your predicted wall time was within 20% of measured, your mental model of virtual-thread scheduling is solid; if it was far off, re-read section 3 and figure out which cost you forgot.

Comments

Popular posts from this blog

JSP Servlet Interview Questions For Freshers Series 1

Java Banking Finance Services and Insurance (BFSI) domain interview questions

Java program to check even or odd number