Caching OpenAI API Responses in Node.js
Caching OpenAI API responses is the fix most teams skip until their first real bill shows up, because the same handful of questions get asked constantly in production — a support bot answering "how do I reset my password" for the hundredth time, a product search endpoint re-describing the same SKU — and every single one of those calls goes all the way out to the model instead of coming back from memory. I added a caching layer to a client's AI support widget a few weeks ago and it cut the OpenAI bill by close to 40% in the first week, without touching a single prompt.
The part that actually takes judgment isn't whether to cache — it's which caching strategy to use, and the two options behave very differently once real traffic hits them.
Two Caching Strategies, Not One
Exact-match caching hashes the request (model, prompt, temperature, and any other parameters that affect the output) and looks up that exact hash in Redis before calling the API. It's a cache hit only when the input is byte-for-byte identical to something you've seen before.
Semantic caching goes further: it embeds the incoming prompt, searches for a previously-cached prompt with a similar embedding above some similarity threshold, and serves that cached response even though the wording doesn't match exactly. "How do I reset my password" and "I forgot my password, how do I change it" would hit the same semantic cache entry but miss an exact-match one entirely.
| Exact-match caching | Semantic caching | |
|---|---|---|
| Setup complexity | Low — hash + Redis GET/SET | Higher — needs an embedding model + vector search |
| Extra latency per request | Negligible | One embedding call on a miss (small, but not zero) |
| Extra cost per request | None | Small embedding API cost on every request |
| Cache hit rate on varied phrasing | Low | Much higher |
| Risk of a wrong/stale answer | Very low | Real, if the similarity threshold is too loose |
| Good first step for most apps | Yes | Only once exact-match alone isn't enough |
Why I Default to Exact-Match First
Most teams jump straight to semantic caching because it sounds more impressive, and I think that's backwards for a first pass. Exact-match caching is almost free to add, has zero risk of returning a wrong answer for a question that only sounds similar, and already covers more traffic than people expect — autocomplete suggestions, repeated form validations, anything your frontend fires on a timer or a keystroke debounce tends to send the exact same payload over and over.
Semantic caching earns its complexity once you've measured your exact-match hit rate and it's genuinely low — a support bot where users phrase the same question a dozen different ways is a legitimate case for it. But I've seen teams reach for semantic caching on day one, tune the similarity threshold too loosely trying to boost the hit rate, and end up serving a cached answer for a question that was actually different enough to need a fresh response. That's a worse failure mode than a cache miss — a cache miss just costs you an API call, a bad semantic match costs you user trust.
Implementing Exact-Match Caching in Node.js
Here's the version I actually ship — it wraps the OpenAI call, hashes the relevant parameters, and checks Redis before doing anything else:
const crypto = require("crypto");
const Redis = require("ioredis");
const OpenAI = require("openai");
const redis = new Redis(process.env.REDIS_URL);
const client = new OpenAI();
const TTL_SECONDS = 60 * 60 * 24; // 24 hours
function cacheKey(params) {
const relevant = {
model: params.model,
messages: params.messages,
temperature: params.temperature ?? 1,
};
const hash = crypto
.createHash("sha256")
.update(JSON.stringify(relevant))
.digest("hex");
return `openai:cache:${hash}`;
}
async function cachedCompletion(params) {
const key = cacheKey(params);
const cached = await redis.get(key);
if (cached) return JSON.parse(cached);
const response = await client.chat.completions.create(params);
await redis.set(key, JSON.stringify(response), "EX", TTL_SECONDS);
return response;
}
A couple of things that matter here: the temperature has to be part of the hash, since a temperature above 0 means the same prompt can legitimately produce different outputs and caching it anyway would hide that variability. And the TTL isn't optional — I've seen a cached answer about pricing or a promo go stale and keep serving for days because nobody set an expiry.
When Caching Can Actually Hurt You
Caching isn't free of tradeoffs just because it saves money.
- Anything personalized (a response that references the user's name, account state, or recent activity) shouldn't be cached by the raw prompt alone — you'll leak one user's cached response to another unless the cache key includes a user or session identifier.
- Time-sensitive answers ("what's today's date," "what's the current price") will cache a wrong answer just as happily as a right one.
- A short TTL matters more than a long one for anything that references data that changes — I'd rather take the extra API cost than serve a stale answer for 24 hours.
Frequently Asked Questions
Will caching make my responses feel less "alive" or personalized? Only if you cache things that genuinely vary per user. Keep personalized or time-sensitive prompts out of the cache entirely, or key them by user ID with a short TTL, and this isn't a real concern.
Does Redis need to be dedicated just for this, or can I reuse what I already have? Reuse what you already have. A caching layer like this adds a small, predictable amount of key volume — it doesn't need its own cluster unless your traffic is already large enough that you'd be scaling Redis anyway.
What similarity threshold should I use for semantic caching? There's no universal number — it depends on your embedding model and how risk-tolerant your use case is. Start high (strict) around 0.95 cosine similarity, measure how often it actually fires, and only loosen it if you're confident a near-miss still deserves the same answer.
Should I cache streaming responses the same way? Not the same way. You'd need to buffer the full stream before caching it, then replay it as a simulated stream on a hit — doable, but it's a separate implementation from caching a plain completion.
Conclusion
Start with exact-match caching — it's a day of work, it's safe, and it quietly cuts real cost on the repeated requests every production app accumulates. Move to semantic caching only after you've measured that your exact-match hit rate is genuinely low and the extra embedding cost and complexity are worth it for your specific traffic pattern, not because it sounds more sophisticated.
0 Comments
No comments yet — be the first to share your thoughts.