Analyzing & Troubleshooting Java Network Lag Using JFR 

Java Flight Recorder (JFR) provides continuous, low-overhead profiling by capturing runtime execution samples directly within the JVM. In this blog, we will examine Network Lag, a performance issue commonly caused by delayed responses from a downstream service or network hop, causing threads to spend excessive time waiting on socket I/O instead of making progress. We will simulate the problem, capture the relevant JFR data, and analyze the recording to identify the underlying issue. Let’s take a closer look.

What are Network Lag?

Fig: Network latency keeps application threads waiting, increasing application response time 

Network lag occurs when calls over the network, whether to a database, a cache, or another service, take noticeably longer to complete than expected. Since a thread making a network call typically waits for the response before continuing, added delay on that call directly adds delay to the thread’s own execution. When many threads are affected at once, this can slow the entire application down, even though the JVM itself isn’t doing anything wrong; the time is being lost waiting on the network.

What causes ‘Network Lag’?

Let’s look at the list of causes for this Network Lag issue: 

  1. Slow or Congested Downstream Services: If a database, cache, or external API the application depends on is responding slowly, every thread waiting on that response inherits the delay.
  2. Network Infrastructure Issues: Problems at the network level, congested links, routing issues, or DNS resolution delays, can add latency to every call passing through, regardless of how well the application code itself is written.
  3. Delayed Downstream Responses: If a downstream service takes longer than expected to respond, whether due to an actual slowdown or a deliberate delay built into a test scenario, every client thread waiting on that response inherits the same delay.

Simulating Network Lag Performance Issue

To understand how Network Lag appears in JFR data, let’s reproduce the issue using a sample Java application.

The following program deliberately runs a downstream server that sleeps for a fixed delay before responding to every request, and a client that repeatedly calls this server and measures how long each call takes.

public static void start() {
NetworkLagServer server = new NetworkLagServer(PORT, INDUCED_DELAY_MS);
Thread serverThread = new Thread(server, "NetworkLagServer");
serverThread.setDaemon(true);
serverThread.start();
sleepQuietly(500);
NetworkLagService service = new NetworkLagService();
long requestNumber = 1;
while (!Thread.currentThread().isInterrupted()) {
try {
long start = System.currentTimeMillis();
service.fetchFromDownstream("localhost", PORT);
long elapsed = System.currentTimeMillis() - start;
System.out.printf("Request %d completed in %d ms%n", requestNumber, elapsed);
} catch (Exception e) {
System.err.println("Request " + requestNumber + " failed: " + e.getMessage());
}
requestNumber++;
sleepQuietly(REQUEST_INTERVAL_MS);
}
}

final class NetworkLagServer implements Runnable {
public void run() {
try (ServerSocket socket = new ServerSocket(port)) {
while (running) {
try (Socket client = socket.accept()) {
Thread.sleep(delayMs);
client.getOutputStream().write(1);
client.getOutputStream().flush();
}
}
}
}
}

final class NetworkLagService {
void fetchFromDownstream(String host, int port) throws IOException {
try (Socket socket = new Socket(host, port)) {
if (socket.getInputStream().read() == -1) {
throw new IOException("Downstream closed the connection without responding");
}
}
}
}

In this program, NetworkLagServer accepts a connection, then sleeps for INDUCED_DELAY_MS (4 seconds) before writing a single response byte back to the client. NetworkLagService.fetchFromDownstream() opens a socket to this server and blocks on socket.getInputStream().read(), waiting for that response byte. Since the server deliberately sleeps before responding, every call to fetchFromDownstream() takes at least INDUCED_DELAY_MS, and start() repeats this call every REQUEST_INTERVAL_MS (2 seconds), logging how long each request actually took. As the loop continues, every request incurs the same induced delay, and the client thread spends the majority of its time blocked on socket I/O rather than doing any real work.

Capturing JFR Data for Troubleshooting Network Lag

To capture JFR data for troubleshooting the Network Lag issue, we recommend that you follow the steps below:

If you are interested in learning about the other methods available to capture JFR recordings, we recommend that you read ‘How to Capture Java Flight Recorder (JFR)?’ blog.

Step 1: Start a JFR recording against the running JVM:

jcmd {PID} JFR.start \
name=loadTestCapture \
settings=profile \
filename=/tmp/tomcat.jfr

Note: Replace <PID> with the Process ID (PID) of your Java application running inside the container. If you don’t know how to find it, refer to our guide on finding the Java application Process ID (PID) for step-by-step instructions.

Step 2: Let it run while the Network Lag is occurring, then stop it manually: 

jcmd {PID} JFR.stop \
name=loadTestCapture

Step 3: Alternatively, for Fixed-Duration Capture,  you can start a recording that automatically exits after a predefined duration (for example, 15 minutes) in a single step:

jcmd {PID} JFR.start \
name=loadTestCapture \
settings=profile \
duration=15m \
filename=/tmp/tomcat.jfr

Analyzing Network Lag Using the JFR Data

You can analyze JFR recording using yCrash JFRPlayer by following the steps mentioned below.  

Step 1: Install yCrash JFRPlayer, which is available in both cloud and on-premises versions. Use one of the options below to get started:

  • Cloud service: Register and upload your JFR file online.
  • On-Premises: Install and run yCrash JFRPlayer on your local machine or within your organization’s environment.

Step 2: Upload the jfr file to your yCrash JFRPlayer. Once JFR file is uploaded, JFRPlayer parses the JFR file and generates an incident report instantly. 

Fig: Uploading a standalone JFR file to yCrash JFRPlayer 

Step 3: The AI Overview mentions network I/O, but only as a side effect of JFR’s own internal processing, not as the actual root cause. It never identifies the real issue: a downstream service deliberately delaying its response. The recommended action even suggests reducing JFR’s network monitoring scope, the opposite of what’s actually needed here.

Fig: yCrash JFRPlayer AI Overview attributing network I/O to JFR’s own overhead rather than identifying the actual downstream delay

Step 4: “Issues in the Device” is where the real signal shows up: long socket read pauses, the longest at 4.010 seconds and averaging 4.003 seconds, closely matching NetworkLagServer‘s 4 second induced delay, totaling 6 minutes 40 seconds across the recording. 

Fig: yCrash JFRPlayer flagging long socket read pauses averaging 4.003 seconds, matching the induced delay in the source code

Step 5: The Application Network Report confirms this is isolated to a single connection, one host, one established TCP/IP connection, ruling out broader network congestion and pointing specifically to this one socket relationship as the source of the delay.  

Fig: yCrash JFRPlayer’s Network Report showing a single host and one established TCP/IP connection, consistent with the isolated socket delay

Simple, right? Now that we have analyzed the data, we are equipped with all the information to fix the problem.

If you face any challenges while analyzing a JFR file using yCrash JFRPlayer, check out our FAQ for answers to common questions and troubleshooting guidance.

How to fix Network Lag

The following are potential solutions to fix this issue:

  1. Confirm Whether the Delay Is Injected or Genuine: Before treating this as a production issue, confirm whether the downstream service is a testing or chaos engineering component intentionally delaying responses, as it is here, rather than a genuinely slow dependency.
  2. Investigate the Actual Downstream Service: If the lag is genuine, examine the downstream service’s own performance, its CPU, GC, or I/O behavior, rather than the client, since the delay originates there.
  3. Set Appropriate Timeouts: Ensure socket reads have sensible timeouts configured, so a slow or unresponsive downstream service degrades gracefully rather than blocking a thread indefinitely.
  4. Monitor Socket I/O via JFR: A JFR recording showing a thread’s time consistently spent in socket read operations, closely matching a fixed or repeating interval, is a clear signal that a downstream dependency is the source of the delay.

Conclusion

Diagnosing and identifying the root cause of Network Lag can be challenging, especially in complex production environments where multiple symptoms often overlap. In this example, the JFR recording was analyzed using yCrash JFRPlayer, which automatically identified the key performance bottlenecks and correlated JVM events to pinpoint the underlying problem. The analysis traced the issue to NetworkLagService.fetchFromDownstream() repeatedly blocking on socket reads while waiting on NetworkLagServer‘s induced delay. By examining the relevant JFR insights, including socket read duration, thread activity, and network connection states, we were able to identify the root cause and determine the appropriate fix.

Annya Arun
Annya Arun
Articles: 33

Share your Thoughts!

Discover more from yCrash

Subscribe now to keep reading and get access to the full archive.

Continue reading