How to Troubleshoot ReDoS in Node.js

Production issues in Node.js applications can be challenging to diagnose, especially when different problems exhibit similar symptoms. To make troubleshooting easier, we’re introducing a Node.js Production Troubleshooting Series, where each article focuses on one specific production issue, explores its underlying causes, and demonstrates a practical, step-by-step approach to identifying the root cause.

ReDoS is one of those production issues that looks like a random freeze until you open the CPU profile. A regular expression with nested quantifiers can spend seconds backtracking on a crafted input. While that runs on the main thread, timers, I/O, and HTTP responses all wait. Average CPU may look busy, but the real clue is a named hot method next to a long Event Loop Lag spike.

In this blog, we’ll walk through how ReDoS shows up in a Node.js application, what typically causes it, and how to trace it back to the exact root cause. We’ll also simulate the issue in a sample application, capture diagnostic data as it happens, and analyze it step by step to pinpoint what’s going wrong.

What is ‘ReDoS’ in Node.js?

Node.js runs your application on a single main thread. That thread also runs the event loop, so timers, I/O callbacks, and HTTP responses all depend on it. When your code calls `RegExp.test` or `String.match` on a pathological pattern, that work runs to completion on the same thread. There is no separate regex worker that finishes it for you.

ReDoS happens when the engine explores an exponential number of ways to match the input before it finally fails or succeeds. A pattern like `/^(a+)+$/` against a long string of a’s that ends wit00h something else is a classic example. From the outside, the app looks frozen for a second or more. In yCrash, you usually see Event Loop Lag climb into the Fatal range, and in the CPU Profiler a named function such as `redosAttack` sits at the top of Self CPU.

What causes ‘ReDoS’ in Node.js?

Here are the usual list of causes for the ‘ReDoS’ problem.

  1. Nested quantifiers: Patterns like `(a+)+`, `(a*)*`, or `(a|aa)+` that force the engine to try many overlapping matches.
  2. Untrusted input on a hot path: User-controlled strings run through a dangerous regex on every request or timer tick.
  3. Near-match then fail: Inputs that almost satisfy the pattern, then fail at the end, which maximizes backtracking.
  4. Copied patterns without a timeout: Regexes taken from Stack Overflow or old libraries that were never tested against long failing inputs.
  5. Validation on the main thread: Running expensive `.test() / .match()` work inline instead of rejecting bad input early or moving it off the event loop.

How to Simulate ReDoS in Node.js

To understand how ReDoS appears in the diagnostic data, let’s reproduce the issue using a sample Node.js application.

The following program repeatedly calls a named `redosAttack` function. That function runs `/^(a+)+$/.test()` on a string of 25 a’s followed by `’!’`. The input almost matches, then fails, so the engine backtracks hard. Each tick keeps calling `redosAttack` until about 1.2 seconds have passed on the main thread, then yields briefly.

``` JS Code (regex-dos.js)
const EVIL = /^(a+)+$/;
const INPUT = 'a'.repeat(25) + '!'; // near-match then fail — forces backtracking
function redosAttack() {
return EVIL.test(INPUT);
}
function redosAttackUntil(targetMs) {
const end = Date.now() + targetMs;
do {
redosAttack();
} while (Date.now() < end);
}
class ReDoSDemo {
constructor() {
this.isRunning = false;
this.timeoutId = null;
}
_scheduleMainTick() {
if (!this.isRunning) return;
// Keep the main thread in redosAttack until ~1.2s have passed.
redosAttackUntil(1200);
console.log('ReDoS: ~1200ms main-thread stall — capture now');
this.timeoutId = setTimeout(() => this._scheduleMainTick(), 400);
}
start() {
if (this.isRunning) return;
this.isRunning = true;
console.log('Starting ReDoS simulation...');
this._scheduleMainTick();
}
stop() {
clearTimeout(this.timeoutId);
this.isRunning = false;
}
}
module.exports = new ReDoSDemo();
```

How to Capture Diagnostic Data for Troubleshooting ReDoS

yCrash is a diagnostic tool that captures runtime performance data from a live Node.js process, including CPU usage, event loop lag, worker thread activity, and call stacks, and turns that data into a report that points to the root cause. To set up yCrash for your Node.js application, follow the steps below:

1. Install yCrash

yCrash is available in both Cloud and On-premises versions. Use one of the options below to get started:

  • Cloud Service: Create an account to use the cloud version.
  • On-premises: Register for a 14-day trial and install the yCrash application on your local machine or within the organization’s environment.

2. Set up the yCrash Node.js hook first

Please note that this is a one-time setup. The hook must be loaded into your Node.js application before yCrash can capture the Node.js internals data that matters for diagnosis, things like worker thread CPU, event loop lag, method profiling, and call stacks. To set this up, follow the Hook Mode steps here: Node.js Diagnostic Capture. This hook setup adds very minimalistic, almost zero overhead to your application. 

3. Launch yc-360 script

Configure and start the yc-360 script to monitor your Node.js process. The full setup and configuration options specific to your deployment is available here: Micro-metrics Monitoring (M3) Mode

When the yc-360 script detects a problem, it captures 360° diagnostic artifacts (GC log, Event loop lags, Worker threads, CPU Profile, Unhandled rejections, Call stack, Active handles, Process, Storage, Kernel, Network, etc.). It uploads them to the yCrash server automatically. yCrash server analyzes the artifacts using advanced pattern recognition and ML algorithms to identify the root cause of the problem.

How to Analyze a ReDoS Using the Diagnostic Data

In this section, we’ll see how to analyze a ReDoS issue using the diagnostic data captured by yCrash. We’ll start with the incident on the calendar dashboard, check the RCA Summary to understand what’s going on, and then drill into Event Loop Lag and the CPU profile reports to find the regex that is blocking the main thread.

Step 1: Open the incident from the yCrash dashboard

When the yc-360 script detects any performance degradation in the Node.js application, it automatically captures 360° diagnostic data and transmits it to the yCrash server. The yCrash server analyses this data and creates an incident in the yCrash dashboard as shown in the screenshot below. Open that incident from the yCrash calendar dashboard to view the reports.

Fig: yCrash dashboard, incident list view

Step 2: Start with the RCA Summary Page

Before diving into details, let’s start with the RCA Summary. yCrash analyzes the Node.js diagnostic data and brings the key findings together in simple language with their severity levels.

In the RCA Summary, we will focus on two sections:

  • AI Overview
  • Issues in the application and device

AI Overview

At the top of the RCA summary report page, you’ll see an AI Overview section. It gives you an AI-generated summary of the incident, which would make even a non-technical person understand the root cause of the problem. Within the AI Overview section, below the AI-summary, you can also click on the ‘Dive Deeper in AI Mode’ button to open a chat window to ask any question to AI about this incident.

Fig: RCA Summary highlighting the AI Overview


As shown in the AI Overview above, the root cause is a CPU-heavy redosAttack path that blocked the event loop for up to about 1520 ms. The overview also mentions an event listener leak warning. That warning is secondary here; the regex stall is what made the app feel frozen.

Issues in application and device

Fig: RCA Summary highlighting the issues in application and device


As shown in the above screenshot, several application-level issues are reported. The 2 that matter most for ReDoS are:

  • Event loop blocked
  • Hot method

The duplicate ms package warning and the event listener leak warning do not explain the freeze. You can set them aside for this investigation.

Step 3: Confirm the stall in Event Loop Lag

Open the Node.js Internals report from the left navigation as shown in the screenshot below. In the top section of this report called “Observations”, you will find all the problems detected in this report. Now let’s click the “detail here” hyperlink present at the end of the Event loop blocked problem statement. This will take you to the Event Loop Lag section.

Fig: Event Loop Lag in Node.js Internals report

As you can see in the above graph, the lag stays near about 1400-1500 ms for most of the window, with Fatal Spike rows such as 1388 ms, 1385 ms, 1389 ms, and a peak of 1520 ms at 16:10:21. That matches the ~1.2 second redosAttack ticks almost one for one.

Event Loop Lag does not name the function. To find out which code burned that time, we will review the CPU Profiling report next.

Step 4: Use CPU Profiling to close the investigation

In the CPU Profiler report page, yCrash aggregates the V8 sampling profile into a Function Explorer table (hot methods ranked by self CPU %). It provides an option to inspect the hot methods to do a deep diagnosis.

Open the CPU Profiler report page from the left navigation and scroll to the Function Explorer section.

Fig: Function Explorer in CPU Profiler report


The Function Explorer confirms what we saw on the RCA Summary page: `redosAttack` used 78.0% Self CPU across about 21,601 calls on one thread. That is the named regex path from the sample app.

For deep diagnosis, let’s click on the “Inspect” button from the same row.

Fig: Inspector view of redosAttack


As you can see in the Inspect view above, `redosAttack` sits at the top of the parent chain. Under it, you can see `redosAttackUntil` and `_scheduleMainTick` in `regex-dos.js`, then Node’s timer path. That closes the loop: the event loop was blocked because the main thread was stuck inside catastrophic regex backtracking, not because it was waiting on I/O.

How to fix ReDoS in Node.js

The following are the potential solutions for this issue:

  1. Rewrite or replace nested-quantifier patterns; prefer simpler, linear-time regexes or a dedicated parser.
  2. Reject or truncate untrusted input before it reaches an expensive `.test()/.match()` call.
  3. Move unavoidable heavy matching off the main thread to a bounded worker, or use a regex engine with timeouts where that is available.
  4. Name hot validation functions so the CPU Profiler is easy to read the next time this shows up.
  5. Watch Event Loop Lag in production so a new bad pattern fails early instead of only showing up as user-facing freezes.

Conclusion

Diagnosing and identifying the root cause of ReDoS can be challenging when the process is still alive, and the only clue is a busy CPU chart. In this example, the diagnostic data was captured using the yc-360 script and analyzed using yCrash server, which helped correlate Fatal Event Loop Lag with the redosAttack hot method.

The analysis traced the issue to catastrophic backtracking inside redosAttack in regex-dos.js. By examining the RCA Summary, the Event Loop Lag chart, and the CPU Profiler Inspect view, we could connect the Fatal 1520 ms stall to that path.

Simplifying the regex, validating input earlier, or moving heavy matching off the main thread addresses the underlying problem and helps prevent the issue from recurring.

Mahesh Devda
Mahesh Devda
Articles: 5

Share your Thoughts!

Discover more from yCrash

Subscribe now to keep reading and get access to the full archive.

Continue reading