How to Troubleshoot Memory Leaks in Node.js

Production issues in Node.js applications can be challenging to diagnose, especially when different problems exhibit similar symptoms. To make troubleshooting easier, we’re introducing a Node.js Production Troubleshooting Series, where each article focuses on one specific production issue, explores its underlying causes, and demonstrates a practical, step-by-step approach to identifying the root cause.

Memory leak is one of those production issues that seems to come out of nowhere, with heap usage rising, garbage collection runs more often, response times get jittery, and only later does the process sit near the heap limit. By the time it surfaces in APM dashboards or a user reports it, the underlying cause is already buried somewhere in the Node.js runtime, and tracking it down after it has already happened is not that easy.

In this blog, we’ll walk through how a memory leak shows up in a Node.js application, what typically causes it, and how to trace it back to the exact root cause. We’ll also simulate the issue in a sample application, capture diagnostic data as it happens, and analyze it step by step to pinpoint what’s going wrong.

What is ‘Memory Leak’ in Node.js?

A memory leak means your application keeps references to objects that it no longer needs. Because these objects are still reachable, V8 cannot remove them during garbage collection. As this continues, the amount of memory used by your application keeps growing.

A memory leak doesn’t always cause an immediate OutOfMemoryError. The application can continue running for hours or even days while memory slowly keeps increasing.

What causes ‘Memory Leak’ in Node.js?

Here is the usual list of causes for the ‘Memory Leak’ problem.

  1. Global/static caches that keep growing: A cache that keeps adding data with .push() or .set() but never removes old entries can slowly consume more and more memory.
  2. Event listener leaks: Adding listeners with on() or addListener() without removing them later with off() or removeListener() can keep objects in memory longer than necessary. For one-time events, { once: true } can help.
  3. Closures holding on to large objects: A callback can accidentally keep references to large objects even after the request that created them has finished. If the callback stays around for a long time, those objects stay in memory too.
  4. Uncleared timers and handles: Timers such as setInterval() and open sockets can keep objects alive. If they are no longer needed but are not cleared or closed, memory usage can keep growing.
  5. Detached DOM-like or graph structures in long-lived services: Any structure that grows with traffic and never shrinks.

How to Simulate Memory Leak in Node.js

To understand how a memory leak appears in the diagnostic data, let’s reproduce the issue using a sample Node.js application.

The following program keeps creating and retaining unique dense number arrays until the heap usage reaches 80% of its limit, then holds there so the application stops allocating more memory. This gives us enough time to capture the diagnostic data while the memory leak is happening and see how it appears in yCrash.

``` JS Code (memory-leak.js)
const v8 = require('v8');
class MemoryLeakDemo {
constructor() {
this.isRunning = false;
this.intervalId = null;
this.retained = [];
this.allocationCount = 0;
this.HEAP_CAP_RATIO = 0.80;
}
start() {
if (this.isRunning) return;
this.isRunning = true;
this.retained = [];
console.log('Starting memory leak simulation...');
this.intervalId = setInterval(() => {
const limit = v8.getHeapStatistics().heap_size_limit;
const used = process.memoryUsage().heapUsed;
if (used >= limit * this.HEAP_CAP_RATIO) {
console.log('Memory leak: holding heap at cap — capture now');
return;
}
const count = Math.floor((8 * 1024 * 1024) / 8);
const arr = new Array(count);
for (let i = 0; i < count; i++) {
arr[i] = (this.allocationCount * 1000003 + i) >>> 0;
}
this.retained.push({ t: this.allocationCount++, arr });
console.log(
`Memory leak: heap ${Math.round(used / 1024 / 1024)}MB ` +
`(${((used / limit) * 100).toFixed(1)}% of ${Math.round(limit / 1024 / 1024)}MB)`
);
}, 2000);
}
stop() {
clearInterval(this.intervalId);
this.retained = [];
this.isRunning = false;
}
}
module.exports = new MemoryLeakDemo();
```

How to Capture Diagnostic Data for Troubleshooting Memory Leak

yCrash is a diagnostic tool that captures runtime performance data from a live Node.js process, including CPU usage, event loop lag, worker thread activity, and call stacks, and turns that data into a report that points to the root cause. To set up yCrash for your Node.js application, follow the steps below:

1. Install yCrash

yCrash is available in both Cloud and On-premises versions. Use one of the options below to get started:

  • Cloud Service: Create an account to use the cloud version.
  • On-premises: Register for a 14-day trial and install the yCrash application on your local machine or within the organization’s environment.

2. Set up the yCrash Node.js hook first

Please note that this is a one-time setup. The hook must be loaded into your Node.js application before yCrash can capture the Node.js internal data that matters for diagnosis, things like worker thread CPU, event loop lag, method profiling, and call stacks. To set this up, follow the Hook Mode steps here: Node.js Diagnostic Capture. This hook setup adds very minimalistic, almost zero overhead to your application. 

3. Launch yc-360 script

Configure and start the yc-360 script to monitor your Node.js process. The full setup and configuration options specific to your deployment are available here: Micro-metrics Monitoring (M3) Mode

When the yc-360 script detects a problem, it captures  360° diagnostic artifacts (GC log, Event loop lags, Worker threads, CPU Profile, Unhandled rejections, Call stack, Active handles, Process, Storage, Kernel, Network, etc.). It uploads them to the yCrash server automatically. yCrash server analyzes the artifacts using advanced pattern recognition and ML algorithms to identify the root cause of the problem.

How to Analyze a Memory Leak Using the Diagnostic Data

In this section, we’ll see how to analyze a Memory Leak using the diagnostic data captured by yCrash. We’ll start with the incident on the calendar dashboard, read the RCA Summary, then confirm retained heap growth in Process Overview, Heap Space, and GC views.

Step 1: Open the incident from the yCrash dashboard

When the yc-360 script detects any performance degradation in the Node.JS application, it automatically captures 360° diagnostic data and transmits it to the yCrash server. The yCrash server analyses this data and creates an incident in the yCrash dashboard as shown in the screenshot below. Open that incident from the yCrash calendar dashboard to view the reports.

Fig: yCrash dashboard, incident list view

Step 2: Start with the RCA Summary Page

Before diving into details, let’s start with the RCA Summary. yCrash analyzes the Node.js diagnostic data and brings the key findings together in simple language with their severity levels.

In the RCA Summary, we will focus on two sections: 

  • AI Overview 
  • Issues in the application and device

AI Overview

Fig: RCA Summary highlighting the AI Overview

As shown in the AI Overview above, the application has a severe memory leak. The system memory usage is high, and a Node.js heap is almost full. If the memory usage continues to grow, the application is at risk of running into an `OutOfMemoryError`.

Issues in application and device

Fig: RCA Summary highlighting the issues in application and device

As shown in the above screenshot, 4 application-level and 2 device-level issues are reported by the yCrash tool. The 2 that matter most for a Memory Leak are:

  • Garbage Collection
  • High heap usage

Step 3: Check Process Overview Resource Usage & Limits

Open the Node.js Internals report from the left navigation as shown in the screenshot below. Then scroll to Process Overview section and select the Resource Usage & Limits tab.

Fig: Process Overview Resource Usage & Limits

In this capture, the numbers here clearly show that the heap is almost full. heap.usedMemory is around 12.09 GB, while heap.totalMemory is about 12.12 GB. The heap.memoryLimit is 14.05 GB

The heap.large_object_space.memorySize alone is using around 12.05 GB. This is where we can see the retained arrays from our application. They are still in memory while the process continues to run.

Step 4: Confirm retained memory in Heap Space

Next, open the Heap Space report from the left navigation as shown in the screenshot below.

Fig: Heap Space report showing high heap usage


The observation section at the top of this report shows the heap is heavily used, at about 86.1% of its 14.05 GB limit. 

The large_object_space is using around 12.05 GB, which matches the large arrays retained by our simulation. The used-vs-available chart is almost entirely consumed the available heap. This confirms that the large amount of memory is still being retained.

Step 5: Confirm the memory leak

Next, open the Garbage Collection report from the left navigation as shown in the screenshot below. At the top of the report, yCrash flags a memory leak in the GC logs.

Fig: GC report detected memory leak


Scroll to the Interactive Graphs section and look at the Heap Usage (after GC) graph.

Fig: GC heap usage after GC


The red triangles in the graph mark the full GC events. Even after garbage collection runs, the heap usage continues to stay high and keeps growing.

This tells us that GC is not able to reclaim the memory because the objects are still being referenced by the application.

When the heap remains high even after GC, it is a strong indication of a memory leak. In our example, this matches the Memory Leak finding in the top section of the Garbage collection report and confirms that the application is retaining memory that it no longer needs.

How to Fix a Memory Leak in Node.js

The following are the potential solutions for this issue:

  1. Use heap dumps to identify what is keeping the objects. Then remove the unnecessary references or put a limit on data that keeps growing.
  2. Make sure event listeners are removed when they are no longer needed. You can use `removeListener()/off()`, `{ once: true }`, or `AbortSignal` where appropriate.
  3. Add cache eviction and maximum sizes. Avoid keeping request history or other data indefinitely.
  4. Increasing `–max-old-space-size` can give the application more memory, but it does not fix the leak. Treat it as a capacity adjustment only after fixing the underlying problem.

Conclusion

Memory Leak diagnosis is about finding out what is keeping the memory alive. 

In this example, yc-360 captured the period when memory usage was high, and yCrash correlated the GC leak detection with Heap Space usage above 86%.

The RCA Summary highlighted the memory leak and high heap usage. The Heap Space report then showed large_object_space holding the retained data.

The fix is to remove unbounded retention and clean up listeners or handle leaks. Once those references are released, memory can be reclaimed by the garbage collector. It will prevent the heap from continuously growing and causing an OutOfMemoryError.

Mahesh Devda
Mahesh Devda
Articles: 5

Share your Thoughts!

Discover more from yCrash

Subscribe now to keep reading and get access to the full archive.

Continue reading