Debug a Node.js Memory Leak with AI

Published · 1 views
Debug a Node.js Memory Leak with AI

Heap climbs from 80MB to 1.4GB over six hours. No crash, no error, just a server that gets slower until you restart it. That's a Node.js memory leak, and this is the AI prompt that actually finds one instead of just restating what --inspect already told you.

The Symptom Nobody Can Grep Their Way Out Of

A memory leak in Node.js doesn't throw. It just sits there, growing, until a pod gets OOMKilled at 3am and someone pastes a heap snapshot into Slack asking if anyone's free. I had this exact situation on a client's order-processing worker last year — steady traffic, no spikes, heap still crept up linearly for days. console.logging object counts tells you that something is growing, never why, because the growth usually isn't in your business logic. It's in something holding a reference it shouldn't.

Chrome DevTools' heap snapshot diff view will show you retained object counts climbing. It will not tell you which closure is the culprit, because that requires reading a retainer tree that's often 40 frames deep across your code, three npm packages, and the event loop itself. That's the part worth handing to an LLM — not instead of the heap snapshot, but on top of it.

Why "Just Use --inspect" Isn't the Whole Answer

Everyone's first answer here is node --inspect plus Chrome DevTools, and that part is correct — you genuinely need the snapshot. The part people skip is that a raw snapshot is thousands of objects with retainer paths like (closure) → arguments → Object → WeakMap → (closure), and manually tracing which of those paths is your bug versus normal Node.js internals takes real time. I've spent entire afternoons doing exactly that by hand.

What actually speeds this up is exporting the retainer paths for the objects with suspiciously high "Retained Size" and feeding those paths to an AI model alongside the relevant source file. The model isn't debugging blind — it's pattern-matching retainer chains against known leak shapes (closures over request-scoped data, forgotten setInterval/setTimeout handles, event listeners added without .removeListener, unbounded caches) far faster than a human reads the same tree.

The Actual Prompt

This is the prompt structure that's worked for me across a few different leak investigations, not a generic "help me debug this" ask:

I have a Node.js memory leak. Heap grows ~15MB/hour under steady
traffic and never drops after GC.

Here's the retainer path for the top offending object from a
Chrome DevTools heap snapshot (Retained Size: 340MB):

(closure in processOrder) -> context -> orderCache (Map, 48,200 entries)
  -> [value] -> Object (order record, ~7KB avg)

Relevant source (orderWorker.js):
<paste the file or the function in question>

Questions:
1. What specifically in this retainer chain indicates the leak?
2. Is `orderCache` unbounded, and if so, where does it get cleared?
3. Give me the minimal code change that fixes this without
   changing the function's external behavior.

Pasting the retainer path, not just the symptom, is what makes this useful instead of generic. A model told "my heap keeps growing" will give you a checklist of common leak causes. A model given the actual retainer chain can point at orderCache by name and tell you it's a Map that's never evicted — which, in my case, it was.

Reading the Answer Without Trusting It Blindly

The AI will usually nail the diagnosis when you give it a real retainer path — that's pattern matching it's genuinely good at. Where I'd push back on its output every time: the suggested fix. It'll often propose wrapping the cache in a library (lru-cache, a TTL map) as the first suggestion, which is sometimes overkill for a problem that's really "we never call .delete() on an order once it ships." Read the proposed diff like you'd read a junior dev's PR — the diagnosis buys trust, the fix still needs your judgment about what this specific service actually needs.

The Real Fix: An Unbounded Cache, Not a Framework Problem

In the case above, the fix had nothing to do with Node.js internals and everything to do with application logic:

// Leak: orderCache never evicts completed orders
const orderCache = new Map();

async function processOrder(order) {
  orderCache.set(order.id, order);
  await shipOrder(order);
  // nothing ever deletes the entry — orderCache grows forever
}

// Fix: evict once the order's lifecycle actually ends
const orderCache = new Map();

async function processOrder(order) {
  orderCache.set(order.id, order);
  try {
    await shipOrder(order);
  } finally {
    orderCache.delete(order.id);
  }
}

No new dependency, no LRU library, just a finally block that matches the cache's actual lifetime to the order's actual lifetime. That's usually how these resolve — the AI's retainer-chain read gets you to "it's this Map," and the fix is whatever your domain logic actually calls for.

Where This Breaks Down

AI-assisted leak hunting falls apart fast once the leak is spread across multiple modules or involves native addons — the model can't see what it isn't shown, and a 40-frame retainer path across code you didn't paste just gets guessed at. It also won't catch leaks caused by long-lived subscriptions to external services (a Redis pub/sub listener that never unsubscribes, say) unless you specifically point it at that code. Treat this as a fast first pass on the retainer tree, not a replacement for actually understanding your own service's lifecycle.

Frequently Asked Questions

Does this work with the free heap snapshot you get from --inspect or do I need Clinic.js / 0x? The built-in Chrome DevTools snapshot from node --inspect is enough — you're exporting retainer paths as text, which any heap snapshot viewer can show you. Clinic.js and 0x add flame graphs that are useful for CPU profiling, but for a memory leak specifically, the Chrome DevTools heap comparison view is the one that matters here.

Can I just paste the entire heap snapshot file instead of specific retainer paths? You can, but don't — raw .heapsnapshot files are JSON blobs that can run into hundreds of megabytes and blow past any model's context window long before it's useful. Export the top 5-10 offending objects' retainer paths as text from the DevTools UI instead; that's a few hundred lines, not gigabytes.

Key Takeaway

The AI prompt earns its keep specifically because it reads retainer-chain output faster than you can trace it by hand — not because it replaces the heap snapshot step. Export the real retainer path, paste the actual source file alongside it, and treat the suggested fix as a draft PR to review rather than a patch to apply straight.

#nodejs #ai-prompts #debugging #memory-leak #chrome-devtools
Aliyan Faisal

Written by

Aliyan Faisal

Full-stack developer and AI/LLM systems engineer. I build LLM integrations, RAG pipelines and automations, and the web apps and servers behind them.

0 Comments

No comments yet — be the first to share your thoughts.

Leave a comment

Never published.