We are having an issue with our node environment running under high load that we have not been able to find the source of.

A little background: we are running a clustered node application using Express for the http framework. Currently, there are 3 boxes with 8 CPU cores on each, and each box is running a cluster of 6 node workers. The setup seems to work great and I've researched all the suggested methodologies such that I believe the setup is solid. We're running node.js 0.8.1 with Express 2.5.11 and XMLHttpRequest 1.4.2.

Here's the issue: We're doing a "dark launch" test of this product (i.e. the browser client code has javascript ajax calls to our APIs in the background, but is not used on the page or shown to the user). After a few minutes running successfully, the system is throwing:

[RangeError: Maximum call stack size exceeded]

We're catching the error with the 'uncaughtException' event in the cluster controller (which starts each worker), however there is no stack trace available at that level. I've done extensive research on this issue and can't seem to find anyone with a similar error. After combing through EVERY line of code in the system, here's what I know:

  • I cannot find any recursion or circular references. (I've read that this error doesn't always mean a recursion problem, but we've checked; we've actually run tests by removing most of the code anyways and it still happens, see below);
  • I've gone down to 1 worker process per box to try and eliminate the cluster as an issue -- the problem still happens;
  • The problem ONLY happens under high load. Our traffic is approx. 1500 pages per second and, during heavy traffic times, can reach 15000 pages per second (we haven't been able to replicate on a dev environment);
  • The timing of the error being caught varies, but is usually within 15 minutes;
  • The error does NOT seem to impact operation! By this, I
Edit
Report