Java Virtual Threads: Pinning, Pools and Performance

Scope: The examples use stable Java 21 APIs as the minimum baseline. Pinning guidance is version-specific: synchronized-block behavior differs between JDK 21–23 and JDK 24+, and the JFR example must be interpreted against the runtime that records it. The timing numbers in the local fixture are chosen delays, not benchmark results.

Java virtual threads make it practical to express a large number of concurrent, mostly waiting tasks with familiar blocking code. They don’t make CPU work inherently faster, and they don’t remove limits imposed by a database, HTTP service, connection pool, or operating system. A useful starting point is an executor that creates one virtual thread per task, then an explicit limit around any scarce downstream resource.

Pinning is another version-sensitive concern. On JDK 21–23, a virtual thread can pin its carrier thread while blocked inside certain synchronized regions. JDK 24 changed this behavior for synchronized blocks and methods. Native methods and foreign-function calls can still pin a virtual thread, so diagnosis should begin with the actual JDK and workload rather than a blanket rule to replace synchronization.

What virtual threads solve

A virtual thread is a thread managed by the Java runtime rather than a thread that permanently occupies one operating-system thread. When a virtual thread blocks on supported operations, the runtime can usually suspend it and use its carrier thread for other work. A carrier is the platform thread that runs a virtual thread when it is scheduled.

This model can suit a service that handles many requests where each request spends much of its time waiting for network responses, database results, or other blocking operations. The code can stay sequential and readable: make a call, wait for its result, then continue. The runtime can manage many such waiting tasks without requiring one platform thread for every task.

That is about concurrency, not a promise of lower latency. A virtual thread cannot make a slow remote service respond sooner. If 500 tasks all wait on a service that can handle only 20 requests at once, the other tasks still have to wait somewhere. Virtual threads can make that waiting cheaper in some workloads, but they don’t increase the service’s capacity.

CPU cost is a separate issue. Work such as compression, image processing, or a tight computation loop needs processor time whether it runs on a virtual thread or a platform thread. Virtual threads are not a general accelerator for CPU-bound code. Oracle describes them as useful for workloads that spend significant time blocked, rather than as a way to speed up computation. Oracle’s virtual-thread guide covers their intended use and runtime behavior.

One other important distinction: a virtual-thread-per-task executor is not a fixed-size pool. It creates a virtual thread for each submitted task. That avoids reusing threads as the main way to limit task concurrency. When a real resource needs a limit, such as a service with a fixed request allowance, limit access to that resource directly.

A minimal per-task executor

The following complete Java file starts a small local HTTP server. Each request to the server waits briefly before returning a response. The client sends several synchronous HTTP requests using a virtual-thread-per-task executor.

Tasks get individual virtual threads, run on carrier threads or wait for I/O, and return results collected through futures.
Virtual threads help represent many blocking tasks. They do not make a remote service respond faster or add CPU capacity.

The server is deliberately local so the example doesn’t depend on an account, API key, or external service. The delay stands in for a blocking dependency. It is illustrative—not a performance test or a claim about a real network service.

Save this as VirtualHttpDemo.java:

import com.sun.net.httpserver.HttpServer;
import java.io.IOException;
import java.net.InetSocketAddress;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ExecutionException;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import java.util.concurrent.TimeUnit;

public class VirtualHttpDemo {
    public static void main(String[] args) throws Exception {
        HttpServer server = HttpServer.create(
                new InetSocketAddress("127.0.0.1", 0), 0);
        server.createContext("/work", exchange -> {
            try {
                Thread.sleep(400);
                byte[] body = "ready".getBytes(java.nio.charset.StandardCharsets.UTF_8);
                exchange.getResponseHeaders().set("Content-Type", "text/plain");
                exchange.sendResponseHeaders(200, body.length);
                exchange.getResponseBody().write(body);
            } catch (InterruptedException e) {
                Thread.currentThread().interrupt();
                byte[] body = "interrupted".getBytes(
                        java.nio.charset.StandardCharsets.UTF_8);
                exchange.sendResponseHeaders(500, body.length);
                exchange.getResponseBody().write(body);
            } finally {
                exchange.close();
            }
        });

        ExecutorService serverWorkers = Executors.newVirtualThreadPerTaskExecutor();
        server.setExecutor(serverWorkers);
        server.start();
        ExecutorService executor = Executors.newVirtualThreadPerTaskExecutor();

        try {
            int port = server.getAddress().getPort();
            URI uri = URI.create("http://127.0.0.1:" + port + "/work");

            try (HttpClient client = HttpClient.newBuilder()
                    .connectTimeout(java.time.Duration.ofSeconds(2)).build()) {
                List<Future<String>> futures = new ArrayList<>();

                for (int i = 0; i < 6; i++) {
                    int requestNumber = i;
                    futures.add(executor.submit(() -> {
                        HttpRequest request = HttpRequest.newBuilder(uri)
                                .timeout(java.time.Duration.ofSeconds(3))
                                .GET()
                                .build();
                        HttpResponse<String> response = client.send(
                                request, HttpResponse.BodyHandlers.ofString());

                        if (response.statusCode() != 200) {
                            throw new IOException("Request " + requestNumber
                                    + " returned status " + response.statusCode());
                        }
                        return "request " + requestNumber + ": "
                                + response.body();
                    }));
                }

                for (Future<String> future : futures) {
                    try {
                        System.out.println(future.get());
                    } catch (ExecutionException e) {
                        throw new RuntimeException("A request task failed",
                                e.getCause());
                    }
                }
            }
        } finally {
            executor.shutdown();
            try {
                if (!executor.awaitTermination(5, TimeUnit.SECONDS)) {
                    executor.shutdownNow();
                }
            } finally {
                server.stop(0);
                serverWorkers.shutdownNow();
            }
        }
    }
}

Compile and run it with a JDK 21 or later installation:

javac VirtualHttpDemo.java
java VirtualHttpDemo

Illustrative output: if all six requests complete successfully, each line will report a request number and ready. The results are printed in submission order because the loop calls get() on each stored future in that order. The requests themselves can finish in a different order. Use the output to check that every task returned a result.

The calls to HttpClient.send are synchronous: each task waits for its HTTP response before returning. The executor gives each submitted task a virtual thread. The calls to Future.get() collect the results and allow task failures to reach the main thread. If a submitted task throws, get() throws ExecutionException; the example unwraps its cause when reporting failure.

The outer finally block requests executor shutdown and stops the local server even if result collection fails. shutdownNow() is used only if the executor has not terminated within the wait. Interruption and cancellation have details that matter in production, so don’t treat a timeout on Future.get() as proof that all underlying work has stopped. A task may be blocked in code that doesn’t respond promptly to interruption.

Platform threads versus virtual threads

For comparison, the same submitted-task pattern can use a bounded platform-thread executor:

ExecutorService executor = Executors.newFixedThreadPool(8);

This is a small excerpt to substitute in the example, not a complete file. The platform-thread executor runs at most eight tasks at once; extra tasks wait in its work queue. By contrast, newVirtualThreadPerTaskExecutor() creates a virtual thread for each submitted task and doesn’t impose that eight-task cap.

QuestionFixed platform-thread executorVirtual-thread-per-task executor
What limits concurrent task execution?The configured number of worker threads.No fixed worker-pool size; tasks are given virtual threads.
Where might extra submitted tasks wait?Often in the executor’s queue once workers are busy.They can wait in their own tasks, but the downstream service can still be overloaded.
What kind of workload may benefit?Workloads where a fixed worker count is useful for resource control or operational reasons.Many blocking tasks, provided other resources are also controlled.
Does it make CPU work faster?No general guarantee.No; virtual threads don’t add processor capacity.

Neither option is automatically right for every service. A bounded platform-thread executor can provide backpressure by limiting how many tasks run at once, although the queue also needs thought: an unbounded queue can grow while tasks wait. A virtual-thread executor makes task creation lightweight in the intended use cases, but by itself it doesn’t say how many database queries or remote requests are acceptable.

To compare the approaches responsibly, record the JDK build, operating system, processor and memory limits, HTTP client configuration, server behavior, request count, and warm-up procedure. Use the same workload and downstream limits for each run. Record throughput and latency distributions as measurements rather than relying on a short demonstration. Don’t publish an expected speedup without actual measurements from the environment that matters.

For background on the difference between concurrent activities and the threads that execute them, see multitasking and multithreading. For executor scheduling context, see scheduling and virtual threads.

Pinning across JDK versions

Pinning means a virtual thread is temporarily tied to its carrier platform thread, so the carrier can’t be freed to run another virtual thread in the usual way while that condition holds. The practical advice about synchronized depends on the JDK version.

Version comparison: synchronized blocking can pin virtual threads on JDK 21 to 23; JDK 24 changes that behavior, while native and foreign-call cases can remain.
The synchronized guidance changed in JDK 24. Do not apply an older blanket locking recommendation without checking the runtime and actual evidence.
Runtime versionWhat to knowPractical response
JDK 21–23A virtual thread can pin its carrier when it blocks while holding a monitor in a synchronized block or method.Investigate long blocking operations inside synchronized regions if the workload shows signs consistent with pinning. Short critical sections that don’t block need not be replaced automatically.
JDK 24 and laterJDK 24 changed synchronized-block pinning behavior. The earlier advice to replace every blocking synchronized region solely to avoid that form of pinning no longer applies in the same way.Use evidence from the actual runtime. Don’t migrate all synchronization to explicit locks just because an older guide recommended it for JDK 21–23.
All versions discussed hereNative methods and foreign-function calls can still pin virtual threads in relevant cases.Inspect the call path and runtime diagnostic evidence if such boundaries are involved.

Oracle’s JDK 24 migration guide documents the change to synchronized-block pinning. Oracle’s JDK 25 guide also describes pinning around native methods and foreign functions and documents the JFR VirtualThreadPinned event. For the runtime-specific details, see: JDK 24 migration changes and Oracle’s virtual-thread guide.

synchronized provides mutual exclusion: it ensures that only one thread at a time holds a particular monitor. It can also make code easier to protect correctly. Changing every synchronized method to a different locking mechanism without understanding its purpose can introduce mistakes or reduce clarity. The useful question is not “Does this code contain synchronized?” but “Does the runtime and workload show a relevant problem at this boundary?” See synchronized methods for background on the construct.

Bound downstream concurrency

Virtual threads and resource limits solve different problems. If a fictional external service permits at most four active calls from your application, use an explicit limit around the call. A semaphore is one option: it represents a number of permits, and a task must acquire a permit before proceeding.

Semaphore gate acquires one of four permits, performs a downstream call, and releases the permit in finally, including on failure.
The semaphore bounds calls through this gate. Its permit-acquisition timeout is separate from the client’s connection and request timeouts.

The following is a complete class showing the pattern in isolation. It doesn’t include a real HTTP client or service credentials. Replace callExternalService() with your project’s actual call and configure its client timeouts separately.

import java.io.IOException;
import java.util.concurrent.Semaphore;
import java.util.concurrent.TimeUnit;

public class DownstreamGate {
    private final Semaphore permits = new Semaphore(4);

    public String callWithLimit() throws IOException, InterruptedException {
        boolean acquired = permits.tryAcquire(2, TimeUnit.SECONDS);
        if (!acquired) {
            throw new IOException("Downstream concurrency limit is busy");
        }

        try {
            return callExternalService();
        } finally {
            permits.release();
        }
    }

    private String callExternalService() throws IOException {
        // Replace this illustrative result with the real client call.
        return "illustrative response";
    }
}

The timed acquisition is a policy choice: if no permit becomes available within two seconds, the method reports overload rather than waiting indefinitely. That does not set a timeout on the external request itself. Configure connection, request, and response timeouts with the client library used by your application. Their exact APIs depend on that library, so this example doesn’t invent a complete client configuration.

The finally block is essential. It releases the permit after success, an exception, or interruption. If code acquires a permit and forgets to release it on one error path, later tasks can wait even though the downstream service has recovered. Keep the permit around exactly the operation whose concurrency you intend to bound.

A semaphore is not the only option. A connection pool may already limit database connections, for example, while an HTTP client may have its own connection or request limits. Check whether a library’s setting controls connections, in-flight requests, or another unit; those are not necessarily equivalent. Avoid adding a second limit without understanding how it interacts with the first.

There are at least three separate controls to consider:

  • Task creation: how many request or background tasks the application accepts.
  • Downstream concurrency: how many operations may be in flight to a particular service or database.
  • Waiting and timeout policy: how long callers wait for a task, a connection, or a remote response.

Increasing task concurrency without setting sensible bounds can move the bottleneck rather than remove it. A queue may grow, a remote service may reject requests, or callers may spend too long waiting. Choose overload behavior—wait, reject, retry under controlled rules, or return a suitable application error—based on the service contract, not thread count alone.

Diagnostics and failure handling

Java Flight Recorder (JFR) can record runtime events for later inspection. Oracle documents a VirtualThreadPinned event in its JDK 25 guide. A recording command can be started from the command line with the JDK’s jcmd tool:

jcmd <pid> JFR.start name=virtual-threads settings=profile duration=60s filename=virtual-threads.jfr

Replace <pid> with the Java process ID. This is a command example, not a report of a recording. Check the jcmd help and JFR event availability for the JDK actually running the application. Event names, availability, recording settings, and interpretation are runtime-specific; a setting that exists on one release may not have the same availability or meaning on another.

Review the recording using an appropriate JFR viewer and look for the documented pinning event and its stack trace when available. Interpret any event in context: a recorded pinning occurrence is evidence to investigate, not proof by itself that virtual threads are the application’s main bottleneck. On JDK 21–23, inspect whether blocking occurs while a monitor is held. On JDK 24+, don’t assume synchronized code is causing the older form of pinning; check native or foreign-call boundaries and other evidence.

Thread dumps can help answer different questions: which tasks are waiting, which stacks are active, and whether many threads appear to be blocked in the same dependency. The exact dump format varies with JDK and tool. Use the runtime’s supported diagnostic options and preserve enough application context to correlate a blocked stack with a request or downstream call.

Separate a slow dependency from pinning. If many tasks wait in a socket read, a remote service or network path may be slow. If callers are stuck waiting for connections, inspect the connection pool and its queue. If tasks wait to enter a synchronized region, inspect the lock holder and the amount of work done while holding the monitor. Pinning is one possible contributor to carrier availability; it is not a synonym for all waiting.

When a task fails, retrieve its result and handle the cause. In the complete HTTP example, Future.get() is wrapped so ExecutionException becomes a failure with the original cause attached. If you add future.get(timeout, unit), a timeout tells you the result wasn’t ready within that time. It doesn’t guarantee that the task, HTTP request, server-side operation, or other underlying work has been cancelled. Define cancellation behavior explicitly and use client-level timeouts and cancellation mechanisms where supported.

Migration pitfalls

Pooling virtual threads

A virtual-thread-per-task executor is designed around a new virtual thread for each submitted task. Putting virtual threads into a fixed-size pool to imitate a platform-thread pool usually mixes up two questions: how tasks are represented, and how many operations a dependency can handle. Keep the per-task executor when suitable, and cap scarce resources separately.

Unbounded submissions

Lightweight task creation doesn’t mean unlimited work is safe. A caller can submit tasks faster than they complete, increasing memory use or pressure on dependencies. Consider admission control, bounded queues where appropriate, and a clear rejection or waiting policy. Measure workload behavior under realistic arrival rates.

Thread-local assumptions

Code often uses ThreadLocal to store request context or mutable state associated with a thread. With a thread-per-task style, check the lifecycle and memory implications of per-thread values, and verify that context is propagated and cleared correctly by the framework you use. Don’t assume a pooled-worker cleanup strategy automatically fits a different thread lifecycle.

CPU-heavy work and lifecycle

Moving a tight CPU loop to virtual threads does not create more CPU capacity. Choose an execution strategy that suits the work and the application’s processor limits. Also close executors and related resources deliberately. The example uses finally for executor shutdown and server cleanup; production services need a lifecycle plan that integrates with their framework’s start and stop hooks.

If your application schedules repeated work, scheduling semantics are a separate topic from whether an individual task uses a virtual thread. See the scheduling and virtual-threads guide for that distinction.

Adoption checklist

Before changing a blocking service, record what you are trying to improve and what must remain bounded. A useful migration is a controlled comparison, not a switch justified by a general claim about threads.

  • Record the exact JDK vendor and version. Apply the JDK 21–23 versus JDK 24+ pinning guidance correctly.
  • Identify whether the main work is blocking I/O, CPU-heavy, or a mixture.
  • Write down limits for downstream requests, connection pools, queues, and request arrival rates.
  • Test success, non-success HTTP status, task exceptions, interruption, and executor shutdown paths.
  • Check that semaphore permits are released after both normal completion and failure.
  • Use JFR and thread dumps on the chosen runtime; label findings as observations from that environment.
  • Compare approaches using the same workload and disclose resource limits and measurement methods.
  • Define success criteria before measuring, such as acceptable latency distribution, resource use, and error rate.

Start with the local delayed-endpoint example, then replace the illustrative call with a real blocking dependency while keeping its concurrency limit explicit. Confirm which JDK is deployed, inspect its JFR and thread-dump evidence, and verify the failure and cleanup paths. Those checks provide a sound basis for deciding whether virtual threads fit the workload—without assuming they remove downstream limits or make every program faster.

Post a Comment

0 Comments