Streaming LLM Responses: SSE, Backpressure, Partial Parsing
Implementing server-sent events for LLM streaming, handling backpressure in Node.js, and parsing partial JSON tokens without corrupting output.

When you stream LLM responses token by token, three problems collide: the transport protocol, the network's willingness to accept data, and the client's ability to consume incomplete JSON. Most implementations nail one or two of these and silently drop bytes under load—or, worse, crash the process with an out-of-memory error. This article walks through how to handle all three with server-sent events, Node.js backpressure, and incremental JSON parsers that don’t corrupt your output.
Why SSE over WebSockets for LLM streaming
Server-sent events (SSE) are unidirectional: data flows from the server to the client, which is exactly what an LLM response stream is. No handshake upgrade, no framing overhead, no bidirectional complexity you won’t use. SSE works over plain HTTP/1.1 or multiplexed HTTP/2 without extra headers. The browser-native EventSource API handles automatic reconnection, last-event-id tracking, and custom event types—all out of the box.
WebSockets add a full-duplex channel that, for a pure response stream, is wasteful. The upgrade handshake adds latency, and you need a library to manage reconnection. Yes, WebSockets let you send cancellation or tool-call requests mid-stream, but you can just as easily use a separate POST endpoint for that. The trade-off is explicit: SSE for one-way streaming plus a small HTTP control API is simpler than maintaining a WebSocket lifecycle that’s half-idle.
The only case where I’d reach for WebSockets instead is when the client must send streaming input (e.g., user voice transcription) alongside the LLM output. For chat-based LLM responses, SSE wins on architecture and tooling.
Implementing an SSE endpoint in Node.js
An SSE endpoint is just an HTTP response with a specific content type and a streaming format. In Node.js with vanilla http or Express:
const http = require('http');
const { Transform } = require('stream');
http.createServer(async (req, res) => {
if (req.url !== '/stream') {
res.writeHead(404);
res.end();
return;
}
res.writeHead(200, {
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
'Connection': 'keep-alive',
});
// Flush headers immediately – some proxies buffer otherwise
res.flushHeaders();
// Simulate an LLM response stream (real code would pipe from OpenAI, etc.)
const tokens = ['He', 'llo', ' ', 'world', '!'];
for (const token of tokens) {
res.write(`data: ${JSON.stringify({ token })}\n\n`);
await new Promise(r => setTimeout(r, 100));
}
res.write('data: [DONE]\n\n');
res.end();
}).listen(3000);The event format is field: value\n\n. Standard fields are data, id, event, and retry. I always include an id field with a monotonically increasing number so the client can request re-sending from a missed event. The retry field (milliseconds) tells the browser how long to wait before reconnecting after a drop.
One gotcha: Express’s default response handling may buffer or compress the stream. Set res.flushHeaders() and disable any compression middleware on the SSE route. If you use compression, the stream becomes a framed blob and the client won’t see partial events.
Backpressure: why it matters and how to handle it
An LLM can generate tokens faster than a network can push them to a slow client. Without backpressure, every res.write() that fails to flush immediately queues data in the internal buffer. Under sustained load, that buffer grows unboundedly and eventually kills your process with an out-of-memory error.
Node.js exposes backpressure through the return value of res.write(). It returns false when the internal buffer is full and the stream’s high-water mark has been reached. When that happens, you must wait for a 'drain' event before writing again.
Here’s a naive loop that will leak memory vs. a backpressure-aware version:
Naive (dangerous):
async function streamTokens(res, tokens) {
for (const token of tokens) {
res.write(`data: ${token}\n\n`);
// res.write returns true/false, but we ignore it – buffer grows indefinitely
}
res.end();
}Backpressure-aware:
function createBackpressureStream() {
return new Transform({
objectMode: true,
highWaterMark: 16, // keep at most 16 queued objects
transform(token, encoding, callback) {
const chunk = `data: ${JSON.stringify(token)}\n\n`;
const canContinue = this.push(chunk);
if (!canContinue) {
// Pause upstream; drain will resume
this._readableState.pipes.forEach(pipe => pipe.pause());
this.once('drain', () => {
pipe.resume();
callback();
});
} else {
callback();
}
}
});
}In practice, I use a simpler pattern with res.write()’s return value and the response’s drain event directly inside an async generator. The key is to never ignore false from write.
Here’s a comparison table of the two approaches:
| Strategy | Memory growth | Data integrity | Implementation effort |
|---|---|---|---|
| Naive loop (no backpressure) | Unbounded – OOM under load | No data loss (just memory exhaustion) | Minimal |
drain-based pause/resume |
Bounded by high-water mark | No data loss | Moderate |
Piped Transform stream |
Bounded | No data loss | Low (with framework) |
The piped approach using Node’s built-in backpressure (e.g., readable.pipe(writable)) is the least error-prone. For SSE, you can create a readable token stream and pipe it to res. Just ensure the response is writable and not buffered.
Partial parsing of JSON tokens: the real problem
LLMs don’t emit complete JSON objects at once. They stream tokens like {"key": "val then ue"}. Calling JSON.parse() on each chunk fails until the full value arrives. Parsing only at the end defeats the purpose of streaming—the client waits just as long as if you sent the whole response.
The solution is an incremental JSON parser that yields structured data as soon as enough bytes arrive. A few options:
stream-json: streams JSON tokens (start object, property, string value, etc.). It emits events as it parses, so you can build partial objects from the stream.jsonrepair: can fix truncated JSON, but it’s better for single repairs than real-time streaming.- A manual accumulator with
try-catch: accumulate incoming chunks, attemptJSON.parse, catchSyntaxErrorand wait for more data. Simple, but it fails if the stream ends on an incomplete chunk (e.g., truncated after a comma). You can mitigate by appending dummy closing brackets and using a lenient parser.
I prefer stream-json for production. Here’s a minimal integration:
const { parser } = require('stream-json');
const { streamValues } = require('stream-json/streamers/StreamValues');
// ss is the SSE response stream
llmStream.pipe(parser()).pipe(streamValues()).on('data', ({ value }) => {
// value is a partially built object – it may have all keys, but some values incomplete?
// stream-json only emits complete values (root objects or top-level array items).
// For nested objects, use stream-object or write a custom filter.
res.write(`data: ${JSON.stringify({ partial: value })}\n\n`);
});The pitfall: most incremental parsers emit complete top-level values only. If your LLM produces a single large JSON object, you may not see any parsed output until the closing brace arrives. In that case, you either need a token-level stream (like stream-json's low-level Parser) or a bespoke parser that yields partial updates for specific keys. For simple chat responses where each token is the next word, a simple try-catch accumulator is often enough.
The related problem of ensuring the LLM actually produces valid JSON in the first place is orthogonal, but I’ve covered it in Reliable Structured Outputs from LLMs Using JSON Schema.
Handling client disconnection and resumption
When a client disconnects, the HTTP request emits a 'close' event. If you don’t abort the LLM API call, you’ll burn tokens on a stream nobody reads—and pay for them. Listen on req.on('close', ...) and call abortController.abort():
const abortController = new AbortController();
req.on('close', () => {
if (!res.writableEnded) {
abortController.abort();
res.end();
}
});
const llmStream = await fetchLLMResponse({ signal: abortController.signal });For resumption, SSE’s Last-Event-ID header is the standard mechanism. The client sends it automatically on reconnect, and the server looks up buffered events after that ID. Implement a small ring buffer in memory (configurable size, e.g., 200 of the most recent events). On reconnect, replay events from the buffer starting after the last emitted ID, then continue with the live stream.
A ring buffer avoids storing the entire conversation forever. If the buffer is too small, the client will miss events and may need a full restart. Choose a size that covers typical reconnect delays (a few seconds of streaming). For long-running streams, you can also persist event IDs and replay from a database—but that’s usually overkill for a chat.
For completeness, see Building an agent loop: tool calls, retries, failure modes for handling mid-stream tool calls that require bidirectional messaging.
Gotchas and limits
Proxy timeouts. Many reverse proxies (AWS ALB, Nginx, Cloudflare) have idle timeout defaults of 60–120 seconds. If your LLM stream is silent for that long (e.g., waiting for user input or a tool call), the proxy will close the connection. Send a keepalive comment event: :keepalive\n\n every 30 seconds. SSE comments (lines starting with :) are ignored by the client but keep the connection alive.
Buffer flushing. Express and other frameworks may buffer writes because of middleware. Always call res.flushHeaders() after setting headers, and ensure no compression or body-parsing middleware is applied to the SSE route. If you must compress, use a streaming compressor that flushes after each chunk (e.g., zlib.createGzip({ flush: zlib.constants.Z_SYNC_FLUSH })).
Multi-byte UTF-8 characters. A single character (e.g., emoji) can be split across two chunks. If your SSE driver counts bytes, it may cut a character in half, producing invalid UTF-8. Always use Buffer.byteLength (not string.length) when computing event sizes, and handle partial surrogates. Most LLM SDKs emit complete characters, but it’s still a risk when you do custom buffering.
Chunked transfer encoding vs SSE
Chunked transfer encoding is a lower-level HTTP mechanism that lets the server send data in arbitrary chunks without a Content-Length header. SSE is built on top of chunked encoding but adds a structured format (data: ...\n\n), event IDs, and automatic reconnection logic in the browser EventSource object.
If you control both the client and server, you could use raw chunked streaming to minimize protocol overhead—just send raw JSON chunks. But you lose reconnection, event filtering, and the standard retry mechanism. For a browser client, SSE is the clear choice because EventSource does the hard work of parsing and reconnecting. For a server-to-server stream where you manage the client code, chunked encoding may be sufficient, but you’ll have to implement reconnection yourself.
For most LLM streaming scenarios, SSE is the right abstraction.
Key takeaways
- SSE is simpler than WebSockets for one-directional LLM response streams; use a separate POST endpoint for cancellation or tool calls.
- Always honor
res.write()returningfalseand use the'drain'event to avoid unbounded memory growth under backpressure. - Partial JSON parsing requires incremental parsers like
stream-jsonor a carefultry-catchaccumulator; streaming is pointless if the client waits for the complete JSON. - Handle client disconnection by aborting the LLM request immediately to save tokens, and implement reconnect using SSE’s
Last-Event-IDwith a ring buffer. - Watch for proxy timeouts, middleware buffering, and multi-byte character splits—these are the silent killers of production SSE streams.
Frequently asked questions
- How do I cancel an in-progress LLM stream when the client disconnects?
- Listen to the `close` event on the request object in Node.js, then call `abort()` on the AbortController passed to the LLM API. This prevents wasted tokens and cost.
- Can I use SSE with HTTP/2 or HTTP/3?
- Yes, SSE works over HTTP/2 and HTTP/3. However, some HTTP/2 implementations may buffer frames; you need to flush explicitly. SSE's long-polling fallback is not needed with modern protocols.
- What happens if the client misses some events during a reconnect?
- SSE supports resumption via the `Last-Event-ID` header. On reconnect, the server can check the last received event ID and replay missed events if you stored them. This requires a buffer of recent events on the server.


