Troubleshooting App Unresponsiveness in Jenkins

If you’re part of a top DevOps team, you’re used to high workloads and seemingly impossible deadlines. And somehow, you make it happen: new releases go out on time, and security loopholes get patched as fast as they’re discovered.

Until the day your Jenkins system goes on the blink. You struggle to even log in to launch a build, then wait interminably for the results. Sadly, the deadlines don’t go away.

When Jenkins becomes unreliable, we need to fix it fast.

What can make Jenkins unresponsive? How do we find the root cause, and how do we solve the problem? More importantly, can we prevent it from happening in the first place?

In this article, we’ll explore answers to these questions.

Common Symptoms of an Unresponsive Jenkins System

We’d see different symptoms, depending on whether the problem lies with the Jenkins controller or with a Jenkins agent.

If a single agent is unresponsive, we may see:

  • Agent disconnects;
  • An agent hangs and stops accepting jobs;
  • Build queues grow;
  • Jobs take much longer to complete than usual;
  • The controller may report the agent as being offline, even though the agent is still running;
  • The agent may be killed by the operating system or container management system.

If the controller is unresponsive, we may see:

  • Multiple agents appear to be unresponsive;
  • The UI is slow or possibly entirely frozen;
  • Requests fail to complete;
  • Build scheduling is sluggish, and jobs may fail to start on time;
  • Build queues grow.

Common Causes of Jenkins Unresponsiveness 

If a JVM is unresponsive, it means there’s a bottleneck. Something is blocking normal operations, causing the application to hang. The bottleneck may be in any of several areas, making troubleshooting tricky:

  • Multithreading issues: Locks, deadlocks, or thread pools too small result in excessive wait times;
  • GC issues: heap too small, GC badly configured, object churn, memory leak, or ClassLoader leak. GC events are extremely CPU-hungry, and cause significant system hangs if the GC is overloaded. 
  • CPU saturation: If the device or container has too few CPUs, or rogue processes hog CPU time, critical processes may hang.
  • Slow I/O: Waiting for disk or network resources;
  • Resource exhaustion: Resources such as DB connections, HTTP connections and file handles are finite in number. Processes may hang waiting for a resource to become available.
  • Memory shortage in the device or container: causes excessive page swaps and slows down the entire device.
  • Slow external resources such as database servers or remote APIs cause threads to wait for long periods.
  • Unterminated tight loops: use excessive CPU time, holding up other threads.

In Jenkins, the first step is to establish whether the problem is within the agent or within the controller. The Jenkins controller is a Java process running jenkins.war. Agents may be on the same machine, on a different server or in the cloud. They often run in a container environment such as Docker. Depending on the Jenkins version, the agent process may be named *agent*.jar, *remoting*.jar or *slave*.jar.

Typical causes of an unresponsive Jenkins agent include:

CauseSymptoms
Network between controller and agentController reports agent offline, even though the agent is still running
Build processes: too resource-hungry, or too many build processes (Typically Maven or Gradle)Agent shows as connected, but builds are slow or stalled
Agent host limitations (CPU, memory, disk)Builds fail or hang; agent disappears
DeadlockBuilds fail or hang 
Waiting on external processesBuild process stalls
File system problemsWorkspace operations hang
Not enough executors per agent, or not enough agentsQueues build up; scheduled tasks may not execute
GC issues Agent extremely slow

Typical causes of an unresponsive Jenkins controller include:

CauseSymptoms
Shortage of resources in the device (CPU, memory etc.)UI and build scheduling become sluggish
GC issuesLong pauses; poor response; slow
Builds executing on controllerUI and build scheduling become sluggish; scheduled builds fail to run
Rogue PluginUnresponsive after plugin install/update
Too many builds running complex pipeline scriptsUI and build scheduling become sluggish
Network slow/unreliableAgents repeatedly disconnect; credential checking hangs
Disk issues: slow/unreliable device, excessive IO, Jenkins home too largeFrequently hangs when busy
Too few agents, or too few executors per agentJob queues build up
Excessive agent traffic: too many agents or too “chatty”Hangs during specific jobs
Waiting on external servicesTasks such as SVN access or credential checking hang

How to Diagnose an Unresponsive Jenkins System 

Since so many different things can cause an application to stop responding, we need to take a holistic approach to gathering and analyzing diagnostics. The table below shows the artifacts we need to troubleshoot an unresponsive application. We wouldn’t usually capture a heap dump at the beginning, since the capture process is resource-hungry and may cause an already-struggling application to crash. We only need one if other diagnostics indicate issues in the heap. 

ArtifactRelevant Information
Linux top or Windows Task ManagerProcess running?
 Overall Memory High?
 Overall CPU High?
 JVM Memory High?
 JVM CPU High?
 Which other processes are consuming high CPU or memory?
 Is swap space heavily used?
Network and Disk Usage StatisticsIs the network healthy? Is the disk slow/overloaded?
GC logsPauses: max and average (Latency)
 Throughput
 GC Frequency
 Heap usage patterns
 Object creation rate
3 Thread Dumps taken at intervals of about 10 secondsIs there a deadlock?
 Blocked threads
 What are threads waiting for?
 Stack Trace: Methods not moving. (Check for I/O, DB, Network, remote API, etc., file handles, OS calls)
 Flame graph – Where is most time being spent?
 Thread leaks: Large number of similarly-named threads, often with the same stack trace
 Thread pools: Are they continuously busy?
top -H -p <PID> (or Windows equivalent)Correlate with thread dump to see threads with high CPU usage
Heap DumpOnly if diagnostics indicate memory issues; highlights where excess memory is used.

Thread dumps are especially useful for narrowing down which area of Jenkins is causing the problem, since the stack traces tell us the exact package, class, and instructions that are currently executing. We can find documentation for each Jenkins package and standard plugins in the Jenkins Javadocs. The thread dump tells us the package name; the Javadocs document the package.

Let’s look at a few examples of what we may see in the diagnostics. We used fastThread to analyze the thread dumps, and GCeasy to analyze GC logs.

Jenkins Unresponsive Due to Garbage Collection Issues 

For this example, we deliberately set the heap dump size very low on the Jenkins controller to make sure the GC had to work hard to clear enough memory for new tasks. This was a small test instance, with short builds running intermittently.

We enabled GC logging and allowed Jenkins to run for a short time, monitoring resource usage by thread using top -H -p 1477, where 1477 was the PID of the running Jenkins controller. Most of the time, the highly active threads were the garbage collection workers, as shown in the screen print below.

Fig: Top Command Showing Consistently High GC Activity

This was a good indication that GC wasn’t performing well, so we submitted the GC logs to GCeasy for analysis.

The resulting report included the charts shown below.

Fig: Selection of GCeasy Graphs

Note:

  • The first graph shows heap usage over time, with full GC events indicated by red markers. At certain times, full GCs are running almost continuously, which is a sure sign of GC problems.
  • The heap usage is consistently too close to maximum.
  • The second graph shows that most of GC’s work is causing other threads to pause. This is not healthy. G1GC should carry out most of its work concurrently.
  • The third graph shows key performance indicators. We can see that some GC pauses are very long.

All these diagnostics indicate that in this case, the cause of the problem is GC overload.

Jenkins Unresponsive Due to Locking Issues and Deadlocks 

To prevent critical areas of code from being executed by more than one thread concurrently, developers use synchronization. If synchronization locks aren’t used carefully, they can result in methods waiting for long periods for a lock to be released. In the worst-case scenario, two threads A and B can block each other indefinitely in a deadlock if:

  • A is waiting for Lock X, which is held by B, and
  • B is waiting for Lock Y, which is held by A.

Thread dump analysis immediately shows us exactly where this type of bottleneck exists. Let’s look at a couple of useful fastThread reports: 

Fig: fastThread Report Showing a Deadlock

When we load a dump into fastThread, it immediately highlights the problem if there is a deadlock. From the stack traces displayed, we can find the exact classes and methods affected. Referring the package names back to the Jenkins documentation lets us see which part of Jenkins is affected.

If the lack of response is due to threads waiting on locks, a different section of the fastThread report indicates which thread is holding each lock, and which threads are waiting for it.

Fig: Interactive Blocked Threads Graph

The report allows us to click on any thread to pop up its stack trace, which will pinpoint the area of Jenkins that is affected.

How Stalled Threads Make Jenkins Unresponsive 

Threads may show as runnable in the thread dump, but are waiting on other services, such as disk I/O, network responses, database access, or API calls.

This is where it pays to take three thread dumps and compare them, so we can see which threads are progressing and which are not. We can then look at the stack trace to find out what the stalled threads are waiting for.

fastThread lets us load the three thread dumps and compare them:

Fig: fastThread Comparison

If threads show as runnable on each of the three dumps, there’s a possibility they may be stalled. We can click on the stack trace from each dump to see if they are moving or still in the same place.

Fig: Stack Trace Showing Network-related Task

In the image above, the stack trace shows a network-related task. If it remains on the same instruction across the three dumps, it indicates the network may be slow.

The thread comparison is useful for diagnosing:

  • Network Waits
  • Database Stalls
  • Slow disk I/O
  • Slow API calls

How to Speed Up Jenkins Unresponsiveness Diagnosis 

All of this is very time-consuming. We need to gather a full range of diagnostics and correlate the evidence to find the root cause.

The free open-source tool yc-360 captures a full range of diagnostic artifacts with a single command-line instruction, including all the artifacts we’ve looked at in this article. For information on how to run this utility and a full list of available command-line arguments, see yc-360 Arguments.

To speed things up even further, the yCrash root cause analyzer lets us upload the entire output of yc-360 and puts the information together to come up with AI-detected root causes and comprehensive reports. In many cases, yCrash leads us directly to the bottleneck that’s causing Jenkins to hang, as shown in the image below:

Fig: yCrash Root Cause Analysis

How to Fix Jenkins Unresponsiveness 

The tables below show common causes of bottlenecks in Jenkins controller and agent, together with suggested fixes for each.

How to Fix an Unresponsive Jenkins Controller 

ProblemDiagnostic FindingsSolution
Device is under-resourcedHigh CPU or memory usageAdd more memory or CPU
GC issuesGC logs show low throughput or excessive latencyTune the GC algorithm; increase heap size if necessary
Builds running on controller using too many resourcesProcesses other than Jenkins using resourcesCreate one or more agents, and disable builds on the Jenkins controller
Plugin IssuesProblem began soon after installing or upgrading pluginEnsure the latest version of the plugin is installed, or remove the plugin
Network UnreliableThread dump comparison shows stalled socket, stream input/output or TLS transport-related tasksSee Troubleshooting Networks
Too many builds running complex pipeline scriptsHigh CPU and thread dumps show CPS-related or Groovy-related threads that are continuously busy.Avoid too many small pipeline steps: rather create scripts for actions that should be performed sequentially
Disk issues: slow/unreliable device, excessive IO, Jenkins home too largeToo many threads waiting on I/O related tasks between dumpsCheck size of Jenkins home; follow disk troubleshooting guides for Windows or Linux
Too few agents, or too few executors per agentBuild queues increasingConfigure more executors or more agents in Jenkins
Excessive agent traffic, such as very large logsInstall plugins such as Support Core or OpenTelemetry to get more detailsReduce “chattiness” in logs, etc.
Waiting on external servicesThread dump comparison shows threads stalling on external calls, e.g., database, SVNInvestigate network and external service performance

How to Fix an Unresponsive Jenkins Agent 

ProblemDiagnostic FindingsSolution
Network Issues Between Agent and ControllerThread dump comparison shows stalled socket, stream input/output, or TLS transport tasksSee Troubleshooting Networks
Resource-hungry build processesTop command shows processes other than Jenkins agent consuming resourcesInvestigate and simplify currently-running builds, or add more resources to the container/device
Agent host limitations (CPU, memory, disk)Disk, memory, or CPU usage high, or disk I/O slowAdd more resources
Locking contentionThread dump shows deadlock or lock waits between dumpsSee Troubleshooting Deadlocks in Jenkins and Troubleshooting Blocked Threads in Jenkins
Waiting on external processesThread dump stalled on DB access, REST calls, SVN access, etc.Check network; investigate and fix external process
File system problemsThreads stalled on file accessMove workspace to a reliable fast device with enough space
Not enough executorsQueues building upIncrease number of agents or executors per agent in Jenkins settings
GC issuesGC logs show low throughput or excessive latencyTune the GC algorithm; increase heap size if necessary

How to Prevent Jenkins Unresponsiveness 

Regular monitoring is the best way of preventing system performance outages. 

We suggest implementing a schedule that uses the yc-360 script to gather comprehensive diagnostics. Check:

  • The output from the top command to look for excessive memory and CPU usage;
  • Network statistics to make sure a faulty or under-configured network is not slowing the system;
  • GC logs to see if the heap usage pattern is healthy, and to check whether throughput is decreasing or latency increasing over time;
  • Thread dumps to make sure threads are not being kept waiting for locks or external services.
  • Disk usage diagnostics to pick up slow or over-full devices. 

For large, critical Jenkins instances, we recommend automatic monitoring via yCrash Agent. This agent samples important JVM diagnostics automatically, and if any problems are developing, it raises an alert. This allows us to troubleshoot and fix issues before they affect our workflow.

Conclusion: Diagnosing and Fixing Jenkins Unresponsiveness 

As with any computer system, diagnosing Jenkins unresponsiveness is tricky. So many different issues can cause a system to slow down that finding the real cause can be time-consuming.

We need to take a holistic approach, capturing and analyzing a wide range of diagnostic artifacts.

yCrash can speed up this process considerably, helping us to get our builds back on track without wasting time.

Jill Thornhill
Jill Thornhill
Articles: 10

Share your Thoughts!

Discover more from yCrash

Subscribe now to keep reading and get access to the full archive.

Continue reading