Stream OpenAI API Responses in Node.js with SSE
Your frontend shows a nice typing-cursor animation, but your Node.js backend calls the OpenAI API, waits for the entire completion, and only then sends one giant JSON blob back to the client. The result: users stare at a blank screen for three to eight seconds on longer completions, then the whole answer dumps in at once. If you're building anything that talks to the OpenAI API from a Node.js backend, streaming the response token-by-token with Server-Sent Events (SSE) is what actually fixes this — and it's simpler to implement correctly than most people expect, as long as you avoid a handful of specific traps.
This post walks through streaming OpenAI API responses in Node.js using SSE, end to end: the Express route, the client-side consumer, and the mistakes that silently break streaming in production even when it works fine on localhost.
Why Streaming Matters for AI-Powered Apps
A non-streamed OpenAI request blocks on the full completion before your server can respond. For a 400-token answer, that's routinely 3-10 seconds of dead air, and the perceived latency gets worse as prompts and max_tokens grow. Streaming changes the shape of that wait: instead of one long pause, the client starts rendering text within a few hundred milliseconds and keeps appending chunks as they arrive. Total generation time doesn't change, but perceived latency drops dramatically, which is why every major AI chat product streams by default.
Server-Sent Events is the right transport for this on a typical REST/Express backend — it's plain HTTP, works through most proxies once configured correctly, and doesn't need the full complexity of WebSockets for a one-directional server-to-client stream.
How OpenAI's Streaming API Works
Pass stream: true to the Chat Completions (or Responses) API call, and instead of one JSON object, OpenAI returns a series of Server-Sent Events, each containing a small delta of the response — sometimes a full word, often just a few characters. Your Node.js backend's job is to read each chunk from the OpenAI SDK's async iterator and immediately forward it to the connected client as its own SSE event, rather than buffering it.
Setting Up an Express Endpoint with Server-Sent Events
The core pattern: set SSE headers, open the OpenAI stream, and write each delta to the response as it arrives.
import express from "express";
import OpenAI from "openai";
const app = express();
app.use(express.json());
const openai = new OpenAI();
app.post("/api/chat/stream", async (req, res) => {
const { messages } = req.body;
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
res.setHeader("Connection", "keep-alive");
res.flushHeaders();
const controller = new AbortController();
req.on("close", () => controller.abort());
try {
const stream = await openai.chat.completions.create(
{ model: "gpt-4o-mini", messages, stream: true },
{ signal: controller.signal }
);
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta?.content;
if (delta) {
res.write(`data: ${JSON.stringify({ text: delta })}\n\n`);
}
}
res.write("event: done\ndata: {}\n\n");
} catch (err) {
if (err.name !== "AbortError") {
res.write(`event: error\ndata: ${JSON.stringify({ message: err.message })}\n\n`);
}
} finally {
res.end();
}
});
app.listen(3000);
Three details here matter more than they look: res.flushHeaders() forces the headers out immediately instead of waiting for the first res.write(), req.on("close") wires the client disconnecting to an AbortController so OpenAI stops generating tokens nobody is reading, and every chunk is JSON-encoded on a data: line so the client always gets valid, parseable payloads.
Consuming the Stream on the Frontend
EventSource doesn't support POST bodies or custom headers, so for an authenticated chat endpoint, read the stream manually with fetch and a ReadableStream reader instead:
const res = await fetch("/api/chat/stream", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages }),
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const parts = buffer.split("\n\n");
buffer = parts.pop();
for (const part of parts) {
if (!part.startsWith("data: ")) continue;
const { text } = JSON.parse(part.slice(6));
if (text) appendToUI(text);
}
}
Common Mistakes When Streaming AI Responses
- Nginx buffering the whole response before forwarding it. If you're behind Nginx,
proxy_buffering off;(orX-Accel-Buffering: noas a response header) is required, or Nginx will happily wait for the stream to finish and defeat the entire point — this is the single most common reason streaming works locally but arrives as one lump in production. - Never wiring up disconnect handling. Without
req.on("close")aborting the OpenAI request, a user closing the tab mid-response leaves your server (and your OpenAI bill) generating tokens for a client that's gone. - Forgetting
res.flushHeaders(). Some Node HTTP setups delay sending headers until the first byte of body, which adds visible latency before the client even starts listening. - Treating a mid-stream error as a client-side JSON parse failure. If OpenAI errors out after already streaming a few chunks, send a distinct
event: errorblock rather than letting the connection just drop — the client can then show "response interrupted" instead of silently truncating. - Blocking the event loop with synchronous work inside the
for awaitloop. Heavy synchronous processing per chunk (regex on every token, synchronous logging) adds up across hundreds of chunks and reintroduces the lag you were trying to remove.
Best Practices
- Set a server-side timeout independent of the client so an abandoned OpenAI stream can't run indefinitely on a bug.
- Log token counts after the stream completes (from the final chunk's usage data, or a separate non-streamed count) rather than trying to estimate mid-stream.
- Rate-limit the streaming endpoint the same way you would a normal one — streaming doesn't reduce OpenAI's per-minute request or token limits.
- Keep the SSE payload shape consistent (
{ text },{ event: "done" },{ event: "error" }) so the frontend parser doesn't need special-casing per event type.
Frequently Asked Questions
Does streaming cost more than a normal OpenAI API call? No. Streaming changes how the response is delivered, not how many tokens are generated or billed — a streamed and non-streamed call with identical output costs the same.
Can I use WebSockets instead of SSE for this? Yes, and it's a reasonable choice if you already need bidirectional communication elsewhere in the app. For a straightforward server-to-client token stream, SSE is simpler to implement and debug, and it works over plain HTTP without a separate protocol upgrade.
Why does my stream work on localhost but arrive all at once when deployed?
Almost always a proxy or load balancer buffering the response — check Nginx's proxy_buffering, any CDN in front of the app, and whether a serverless platform's response model supports streaming at all before assuming it's an application bug.
Key Takeaways
Streaming OpenAI API responses in Node.js comes down to three things: forward each chunk from the SDK's async iterator immediately instead of buffering it, tie client disconnects to an AbortController so you're not paying for tokens nobody sees, and check your proxy configuration before debugging application code when a stream arrives instantly in one piece. Get those three right and the rest — the SSE headers, the frontend reader — is boilerplate you write once and reuse across every AI-backed endpoint.
0 Comments
No comments yet — be the first to share your thoughts.