Building ChatGPT-like Streaming AI Chat with SSE: From Zero to Production
This article details the full-stack implementation of SSE-based streaming for an AI assistant, covering protocol design (delta/done/error events), backend challenges (ThreadLocal, Retrofit @Streaming, JSON incremental parsing), frontend fetch/ReadableStream handling, and Nginx deployment fixes for buffering and timeouts.
Last week I rebuilt the entire conversation pipeline of our AI assistant "Ruan Xiao Zhi" from a single blocking response to Server-Sent Events (SSE) streaming, achieving the same character-by-character typewriter effect users see in ChatGPT.
Why SSE
Before coding, I evaluated three common server-to-browser push mechanisms:
Short polling : Client asks "done yet?" every few seconds. Lowest cost but poor real-time feel; suitable for infrequent status refreshes.
WebSocket : Full-duplex long-lived connection with bidirectional messaging. High operational cost — heartbeats, reconnection logic, and state machines are all manual. Best for chat rooms or collaborative editing where the client also pushes continuously.
SSE (Server-Sent Events) : Unidirectional push over plain HTTP. Browser has native EventSource support with automatic reconnection. Server side is just a regular HTTP endpoint. Lowest complexity for one-way streams like AI token output or notifications.
My use case is pure one-way: one question, one answer stream. SSE fits perfectly, so I reserved WebSocket for truly bidirectional scenarios.
Architecture and Event Protocol
The pipeline has three layers:
Frontend (Vue) : Uses native fetch + ReadableStream to consume the stream.
Spring Boot backend : Returns an SseEmitter, offloads the streaming work to a thread pool, and calls the upstream AI gateway.
DeepSeek (OpenAI-compatible gateway) : Returns an SSE data stream line by line.
The backend wraps incremental text into three event types sent over the same connection:
delta — {"text":"incremental text"}: Real-time fragment; frontend renders each piece as it arrives.
done — {"reply":"full reply","links":[...]}: Authoritative final result after the stream ends. Frontend overwrites the streamed text with this payload and appends link cards.
error — {"message":"error hint"}: Fallback for upstream failures or push errors.
Streaming handles the experience; the done event guarantees correctness. Even if the incremental display drifts, the final authoritative result corrects it.
I call this self-healing design: the stream optimizes perceived latency, while done ensures data integrity.
Three Critical Backend Details
1. ThreadLocal Pitfall
Our framework (RuoYi) stores the logged-in user in a ThreadLocal via SecurityUtils. Because the streaming work runs in a thread-pool thread, that ThreadLocal is empty there. Solution: capture user, context, and system prompt in the request thread before submitting to the pool; the async thread only does the streaming push.
@Override public SseEmitter assistantChatStream(String question) {
// Request thread: build user, context, prompt completely
SysUser user = SecurityUtils.getLoginUser().getUser();
AssistantContext ctx = buildAssistantContext();
String prompt = buildAssistantPrompt(user, ctx.getContextText());
SseEmitter emitter = new SseEmitter(180_000L); // 3-minute timeout
// Async thread only does streaming push
threadPoolTaskExecutor.execute(() ->
doStreamChat(emitter, user, question.trim(), prompt, ctx));
return emitter;
}2. Retrofit @Streaming Annotation
DeepSeek exposes an OpenAI-compatible endpoint. Setting stream=true in the request yields an SSE response. However, OkHttp buffers the entire response body by default, defeating streaming. Adding @Streaming to the Retrofit interface method keeps the response body as a stream so we can read line by line.
@Streaming
@POST("api/deepseek/v1/chat/completions")
Observable<ResponseBody> streamChatCompletion(@Body ChatCompletionRequest request);Each upstream line starts with data: and ends with data: [DONE]. We parse each line, extract choices[0].delta.content, accumulate the full content, and feed it to a custom JSON incremental extractor (see below) to emit delta events.
try (BufferedSource source = body.source()) {
String line;
while ((line = source.readUtf8Line()) != null) {
if (!line.startsWith("data:")) continue;
String data = line.substring(5).trim();
if ("[DONE]".equals(data)) break;
JsonNode node = objectMapper.readTree(data);
String content = node.path("choices").path(0)
.path("delta").path("content").asText("");
fullContent.append(content);
// Incrementally extract reply text from JSON fragments
String delta = extractor.feed(content);
if (!delta.isEmpty()) {
emitter.send(SseEmitter.event().name("delta")
.data(toJson(Map.of("text", delta))));
}
}
}3. JSON Incremental Extractor (State Machine)
The AI is instructed to output a JSON object containing a reply string and a links array. But the stream delivers fragments: one chunk may end at {"rep, the next continues ly":"Hello. Waiting for complete JSON would kill the streaming effect. So I built a tiny state machine that pulls the reply string content out of the incomplete JSON on the fly.
The machine has three phases:
Find the reply key : Buffer incoming chunks until the opening quote of the reply value is seen.
Inside the string : Emit each character as it arrives, handling escape sequences ( \n → newline, \uXXXX → Unicode char). Unicode escapes may be split across chunks, so the machine buffers across boundaries.
Closing quote : When an unescaped quote appears, the string ends; subsequent content is discarded.
private String readString(String chunk) {
StringBuilder out = new StringBuilder();
for (int i = 0; i < chunk.length(); i++) {
char c = chunk.charAt(i);
if (escaped) {
escaped = false;
switch (c) {
case 'n': out.append('
'); break;
case 't': out.append('\t'); break;
case 'u': unicodeLeft = 4; break; // cross-chunk assemble
default: out.append(c);
}
continue;
}
if (c == '\\') { escaped = true; continue; }
if (c == '"') { phase = 2; break; } // string end
out.append(c);
}
return out.toString();
}The user only sees plain text; they never know the backend is parsing a growing JSON structure.
Two Frontend Gotchas
1. Axios Cannot Stream
Axios relies on XMLHttpRequest, which does not expose a readable stream. Switch to native fetch and use response.body.getReader() to read chunks incrementally.
2. Packet Fragmentation (Sticky/Partial Packets)
Network frames don't respect SSE event boundaries. One event may be split across two chunks, or two events may arrive in one chunk. The fix: accumulate a buffer, split on blank lines ( \n\n) to isolate complete SSE event blocks, then parse event: and data: lines.
const read = () => {
reader.read().then(({ done, value }) => {
if (done) { finish(() => onError(new Error("AI reply interrupted"))); return; }
buffer += decoder.decode(value, { stream: true });
let idx;
// Split on blank lines to get whole SSE events
while ((idx = buffer.indexOf("
")) >= 0) {
const raw = buffer.slice(0, idx);
buffer = buffer.slice(idx + 2);
if (raw) handleEvent(raw); // dispatch by event name
}
read(); // continue reading
}).catch(err => finish(() => onError(err)));
};Also make finish idempotent: a finished flag ensures done and onError fire at most once, preventing duplicate error toasts when the stream aborts.
Deployment Trap: Nginx Buffering
Locally everything streamed perfectly. On the server, the typewriter effect vanished — the response arrived in one big chunk again. Cause: Nginx's proxy_buffering defaults to on, so it buffers the upstream response before forwarding. SSE needs proxy_buffering off so each flush goes straight to the client. Two ways to disable it:
Add response header X-Accel-Buffering: no in the backend endpoint (one-liner).
Or configure Nginx location block explicitly for SSE paths.
Also increase proxy_read_timeout (default 60s) because long AI generations can take 2–3 minutes. Example Nginx config:
location /api/ {
proxy_pass http://backend;
# SSE critical settings
proxy_buffering off; # disable response buffering
proxy_read_timeout 300s; # longer than max stream duration
proxy_set_header Connection '';
proxy_http_version 1.1;
}After the fix, the first character appears in 1–2 seconds, then streams continuously with a blinking cursor at the tail. Technically it's just text push over one HTTP connection; experientially it's a whole era apart. Users don't care if the model is three or five seconds faster — they care that they can see it working.
The typewriter effect has survived over a century because it gives waiting a shape.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Code Farmer Manor Chronicle
A heart like drifting clouds, ever at ease; a mind like flowing water, free to roam.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
