Troubleshooting Storage Saturation in Jenkins

More and more top organizations worldwide rely on Jenkins to streamline entire project lifecycles. From testing prototypes and automating the CI/CD pipeline through to production workloads, Jenkins provides reliability, speed, and cost savings.

When Jenkins disk space issues occur, the entire organization may be affected. It’s therefore critical to diagnose and fix Jenkins problems quickly.

In this article, we’ll look at causes of disk-related problems, fast troubleshooting techniques, and suggested fixes for common problems. 

What Are the Symptoms of Jenkins Disk Space Issues?

Storage problems cause all kinds of unhealthy symptoms in a Jenkins application.

If the issue affects a Jenkins agent, we may see:

  • Builds slowing down or appearing to hang;
  • Long delays between console outputs;
  • Agents intermittently disconnecting or failing to respond to the controller;
  • Docker container operations becoming painfully slow;
  • Agent queues building up;
  • Agents crashing with disk-related error messages.

If the problem is in the Jenkins controller, the symptoms are different. One or more of these indicate possible disk issues.

  • Builds on multiple agents are slow to start.
  • Delays occur with updating build status and console output.
  • Multiple agents are reported as being offline.
  • Scheduled jobs aren’t run.
  • UI is unresponsive. This is sometimes intermittent.
  • Jenkins restarts are slow.

What Causes Disk Space Problems in Jenkins?

When disk saturation occurs, requests for storage reads or writes can’t be completed on time. Applications waste time waiting for I/O, and performance drops accordingly. This can happen for one of several reasons:

  • A device is more than 90% full;
  • Read/write queues are filling up faster than the hardware can process them. This could be because the hardware is slow, or it could be because too many I/O based processes are running at the same time;
  • Bandwidth problems, such as bottlenecks on ports or links;
  • A disk is faulty;
  • Memory limitations are causing excessive swapping to storage.

The table below shows possible areas where disk space usage in Jenkins controller and agents may be excessive:

Jenkins ControllerJenkins Agent
Excessive build historySource code checkout
Large volume of pipeline build dataCompilation/build
Heavy workspace activityMaven/Gradle dependency downloads
Excessive artifact storageLarge workspaces
High-volume console logsStorage-intensive Tests
Heavy SCM polling/branch indexingLarge Docker images, storage-heavy tasks within Docker or inadequate clean-up
Jenkins backups very large or scheduled during peak hoursVerbose or uncleared Docker container logs
Plugins/plugin dataLarge artifact creation
GC logging: Too verbose, or stored on device with heavy contentionLarge artifact upload
System/log files overly verbose, or stored on high-contention volumesHeavy Git operations
Excessive Temporary filesWorkspace cleanup
Insufficient disk spaceBuild logs
High Swap activityTemporary files
 Disk-full condition
 Swap/memory pressure
 Multiple concurrent builds

Since there are many possible causes, we need to use diagnostic tools to find more information before looking at how to fix the problem. 

How to Diagnose Disk Space and Disk I/O Problems in Jenkins

Let’s now have a look at the tools and tactics we can use to diagnose disk problems, and how we can relate these to Jenkins. We’ll then work through a short case study to illustrate the process.

If we suspect Jenkins may be affected by storage-related problems, we should find the answers to a series of questions to help us establish the root cause.

These are summarized in the table below, along with the tools we can use to find the answers.

QuestionToolExample
Are any devices close to 100% usage?dfdf -h
Which directories use most space?dudu -d1 -h $JENKINS_HOME
Are I/O read/write queues excessive?iostatiostat -x -b /dev/sda1 5
Are there errors in the Jenkins logs?journalctlsudo journalctl -eu jenkins
Are there errors in the kernel logs?dmesgdmesg|grep -I error
Are Jenkins threads stalled on disk-related tasks?fastThreadSee Mastering Thread Dump Analysis 

For a deeper look at solving disk issues in Linux, see Troubleshooting Disk Issues in Linux. For Windows-based systems, see Troubleshooting Performance Problems in Windows.

2. Relating Disk Diagnostics to Jenkins

We need to understand where Jenkins is likely to store information. The du command lets us explore disk usage by directory. The tables below relate this back to Jenkins. 

In the controller, Jenkins usually uses the following directories:

DirectoryMain Use
$JENKINS_HOMEPersistent Jenkins data
$JENKINS_HOME/jobs/Jobs and build records
jobs/…/builds/Build history and artifacts
/var/log/jenkins/Service and JVM logs
workspace/ (Possible)Source and build output
/var/lib/docker/ (Possible)Docker daemon storage
/tmp/Temporary files

The agent is likely to use the following directories:

DirectoryMain Use
/var/log/jenkins/Service and JVM logs
workspace/Source and build output
/var/lib/docker/Docker daemon storage
/tmp/Temporary files
~/.m2/repository/Maven dependencies
~/.gradle/caches/Gradle caches
Agent diagnostic directory (configured in the Docker image)Diagnostic logs

To relate thread dump findings back to Jenkins, we need to understand the Jenkins package structure. By examining the stack traces of threads stalled on disk-related tasks, we can establish which area of Jenkins is experiencing the problem.

The table below summarizes the high-level breakdown of the Jenkins package structure.

Packages
Jenkins Corehudson.model.*
 jenkins.model.*
Jenkins Agenthudson.remoting.*
 org.jenkinsci.remoting.*
Plugins*.plugin,*
3rd Party LibsVarious

For full details of the package structure, see the Jenkins Javadocs.

3. Troubleshooting Jenkins Disk Space Issues: A Practical Example

Let’s see how we would put all that information together in practice.

To simulate disk overload, we set up a test system as follows:

  • We set GC logging to high verbosity and wrote the logs to a USB drive with limited size. To achieve this, we set JAVA_OPTS in the Jenkins configuration as follows:
    • JAVA_OPTS=-Xlog:gc*=trace:file=/media/jill/MyUSB/logs/jenkins.log
  • We created a script designed to keep the USB drive busy by filling it with thousands of unnecessary files:
#!/bin/bash
for i in {1..10000}; do
SUFFIX=$i.$(date +%Y%m%d%H%M%S%N)
cp /home/jill/shenandoah.log /media/jill/MyUSB/jenkins/log$SUFFIX
done
  • We set the number of executors to run on the Jenkins controller to 5, to make sure builds with the label WASTE would run in parallel on the controller
  • We created a Jenkins job to run the above script, directing it to the WASTE label.
  • We enabled Jenkins Thin Backup and directed the output to the USB drive.
  • We requested multiple builds of the above job at the same time as initiating a backup.

Not surprisingly, Jenkins first slowed down, then became totally unresponsive. Let’s see what diagnostics showed us when we worked through the questions suggested earlier.

Are any devices close to 100% usage?

We ran the command df to see disk usage by device:

$ df
Filesystem 1K-blocks Used Available Use% Mounted on
tmpfs 343808 3696 340112 2% /
/dev/sda2 959786032 89375320 821585204 10% /
tmpfs 1719020 0 1719020 0% /dev/shm
tmpfs 5120 4 5116 1% /run/lock
tmpfs 1719020 0 1719020 0% /run/qemu
/dev/sda1 523248 6232 517016 2% /boot/efi
tmpfs 343804 108 343696 1% /run/user/1000
/dev/sdb1 1883644 1796620 0 100% /media/jill/MyUSB

This showed that /dev/sdb1, which was our USB drive mounted on /media/jill/MyUSB, was 100% full. The hard disk, /dev/sda1I, has plenty of space.

Which directories use the most space? 

Once we’d established we were short of space on the USB drive, we used the command 

du -d2 -c -h media/jill 

to see which top-level directories on this drive were using the most space. Note:

  • -d2 specifies we want to explore directories to a depth of 2
  • -c requests an overall total
  • -h displays the sizes in a form that’s easily readable.

Here’s the output:

$ sudo du -d2 -c -h /media/jill
94M /media/jill/MyUSB/logs
314M /media/jill/MyUSB/jenkins
16K /media/jill/MyUSB/lost+found
4.0K /media/jill/MyUSB/backupj
1002M /media/jill/MyUSB/jenkinsb
1.4G /media/jill/MyUSB
1.4G /media/jill
1.4G total

This shows three directories that are using a lot of space:

  • jenkins is the directory used by the script waste_disk
  • jenkinsb is the directory configured for Thin Backup
  • logs is the directory configured for GC logging.

Are I/O read/write queues excessive?

We ran the command iostat -p /dev/sda1,/dev/sdb1 -x 20 to check disk activity levels on the devices /dev/sda1 and /dev/sdb1. 

The text below shows a portion of the output, truncated to fit.

avg-cpu: %user %nice %system %iowait %steal %idle
2.47 0.01 2.00 66.77 0.00 28.75
Device w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_awt dareq-sz f/s f_awt aqu-sz %util
sda1 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
sdb1 28.85 2526.40 68.35 70.32 32.25 87.57 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.93 93.24

We have a high value for IO waits overall (66.77). /dev/sdb1 is heavily utilized, with wait queues and over 90% utilization.

We would therefore conclude that both disk space and high usage are a problem in this application.

Are there errors in the Jenkins logs?

In this application, journalctl showed no errors in the Jenkins log.

Are there errors in the kernel log? 

In this application, dmesg showed no errors in the kernel log.

We captured three thread dumps at 10-second intervals to search for threads that were not progressing between dumps. We then loaded these into fastThread, which analyzed the dumps and created a report.

The image below shows the fastThread summary.

Fig: fastThread Summary

This summary shows threads that are not moving in the following methods:

  • write0() method in sun.nio.ch.UnixFileDispatcherImpl
  • directCopy0() method in sun.nio.fs.LinuxNativeDispatcher

Since both of these are I/O-related, this is highly relevant to our investigation. Clicking on the first link shows us full details of the following thread, including its stack trace:

Fig: Thread Stalled on write() Method

The second link shows this thread:

Fig: Thread Stalled on directCopy() Method

Following the stack traces, we see both are associated with the Thin Backup plugin. This tells us the backup is held up due to disk saturation.

Exploring the fastThread report of threads that are waiting, we see the thread below. The starting point of the stack trace is hudson.model.Executor.run(Executor.java:456). This identifies the thread as one of the executors running a build job. Conveniently, Jenkins names the thread with the name of the job and the build number, in this case Waster #23. The thread dump shows that it is waiting for a PID, which would be a script run by the job.

Fig: Stack Trace for Executor Thread

This makes it easy to tie waiting threads back to actual running jobs.

We’ve established that:

  • Disk space is critically short in /dev/sdb1
  • Writes are queueing for this drive, causing application threads to wait
  • GC logging, Jenkins backups and a Jenkins job are all contending for /dev/sdb1
  • Threads associated with the Thin Backup plugin are stalled on disk-related methods
  • The job Waster is hung, waiting for a script to complete.

With these findings in mind, we can re-look at the list of common causes of Jenkins disk saturation.

How to Fix Disk Space and Disk I/O Problems in Jenkins

Let’s now refer back to the root causes of disk problems in Jenkins, alongside possible fixes by cause.

Within the Jenkins controller:

CauseDiagnostic FindingsSuggested Fix
Build history / build recordsHeavy writes under $JENKINS_HOME/jobs; large or numerous build records; I/O spikes during buildsReduce build retention; use Discard Old Builds; archive only required artifacts
Pipeline build dataLarge build.xml files; extensive Pipeline state under builds/; high I/O from complex pipelinesSimplify pipelines; reduce retained builds; review Pipeline durability settings
Workspace activityLarge or numerous workspaces under $JENKINS_HOME/workspace; controller jobs actually executing buildsMove builds to agents; avoid running builds on the controller
Artifact storageLarge artifacts under jobs/…/builds/…/archive; sustained writes during artifact archivingStore artifacts externally; reduce artifact retention; avoid unnecessarily large artifacts
Console logsLarge log files under build directories; rapid growth during verbose buildsReduce excessive application/build logging; configure retention
SCM polling / branch indexingI/O spikes during SCM scans; many repositories/branches; frequent scansIncrease polling/scan intervals; use webhooks; reduce branch discovery
Jenkins backupsLarge sequential read/write activity during ThinBackup; I/O spike coinciding with backupSchedule backups off-peak; reduce backup scope/retention; write backups to separate storage
Plugins / plugin dataSpecific plugin activity coincides with I/O; files changing outside normal build directoriesIdentify offending plugin; update/configure it; disable if unnecessary
GC logginggc.log grows rapidly; sustained writes while JVM is under GC pressureReduce unnecessary logging in production; rotate logs; put high-volume diagnostic logs on appropriate storage
System/log filesjournalctl/system logs or /var/log/jenkins growing rapidlyIdentify noisy service; configure log rotation; correct underlying errors
Temporary filesHeavy activity under /tmp or Jenkins temporary directoriesIdentify process generating files; clean up; correct workload/plugin causing excessive temporary data
Insufficient disk spacedf shows high utilization; performance deteriorates as filesystem approaches fullRemove/archive unnecessary data; increase filesystem capacity
Swap activityvmstat shows high si/so; disk I/O coincides with memory pressureReduce memory pressure; increase RAM; tune JVM/application memory

Within the Jenkins agent:

CauseDiagnostic FindingsSuggested Fix
Source-code checkoutLarge I/O burst during Git checkout; .git directory grows substantiallyUse shallow/sparse checkout where appropriate; reduce repository size
Compilation/buildHeavy reads/writes in workspace; I/O correlates with compilationUse faster/local storage; optimize build; move expensive work to suitable agents
Maven/Gradle dependency downloadsLarge writes to ~/.m2, Gradle cache, etc.; repeated downloadsPersist dependency caches; use repository mirrors; avoid unnecessary cache deletion
Large workspacesWorkspace grows rapidly; many files; du identifies large directoriesClean workspaces; use deleteDir()/workspace cleanup where appropriate; reduce generated files
TestsHeavy I/O during test execution; large reports, logs, screenshots or temporary filesReduce unnecessary test output; clean test artifacts; separate heavy test workloads
Docker buildsHigh I/O under Docker storage (/var/lib/docker); large image layers/build contextsUse .dockerignore; reduce image/layer size; clean unused images; use BuildKit efficiently
Docker container logsLarge JSON/container log files; rapidly growing *-json.log filesConfigure Docker log rotation; reduce application logging
Artifact creationLarge archives/JARs/WARs generated; I/O spike during packagingAvoid unnecessary packaging; compress selectively; clean intermediate files
Artifact uploadHigh disk reads followed by network transfer; coincides with archiveArtifactsReduce artifact size; archive only required files; use external artifact repository
Git operationsHigh I/O during merge, checkout, branch operations; large .git directoriesShallow/sparse clones; reduce repository size; avoid unnecessary fetches
Workspace cleanupLarge deletion bursts between builds; rm/cleanup dominates I/OClean incrementally; use ephemeral agents; avoid retaining unnecessary files
Build logsLarge local log files or application logs generated by buildsReduce verbosity; rotate logs; redirect unnecessary diagnostic output
Temporary files/tmp or workspace temporary directories rapidly growingIdentify process creating them; clean up; configure application temp directories
Disk-full conditiondf near 100%; builds fail with No space left on device; filesystem latency increasesClean workspace/cache/images; increase disk; enforce quotas/limits
Swap / memory pressureHigh si/so in vmstat; disk activity despite relatively little application I/OAdd memory; reduce concurrent builds; tune JVM/container memory
Multiple concurrent buildsSeveral builds simultaneously reading/writing the same physical disk; high %utilLimit executors; distribute builds across agents; use faster storage

How to Monitor and Prevent Jenkins Disk Space Problems

When we manage critical systems like Jenkins, it’s essential to put proper monitoring and alerts in place. 

Administrators commonly use one or more of the following tools to raise alerts in time to proactively prevent problems.

For Jenkins systems that handle crucial workloads, we recommend automated monitoring with yCrash agent. This low-overhead tool samples many different aspects of a JVM and its environment regularly. If it detects situations that could become problematic, such as disk saturation, it captures a full range of diagnostic information as illustrated in the image below. It submits these to a yCrash server, which uses machine learning to diagnose possible problems and root causes. It raises an alert, and produces a full set of useful diagnostic reports.

Fig: yCrash 360 Degree Diagnostic Artifacts

Conclusion

Jenkins disk space issues play havoc with the tight deadlines DevOps typically have to work with to stay competitive.

It’s essential to be able to identify the root cause and apply the appropriate fix quickly to avoid downtime.

Regular monitoring allows us to pick up developing problems before they affect performance, and correct the problem without affecting critical workflows.

Jill Thornhill
Jill Thornhill
Articles: 10

Share your Thoughts!

Discover more from yCrash

Subscribe now to keep reading and get access to the full archive.

Continue reading