Updated 2026-09-15 (first published 2026-02-19). The argument and the eight layers still hold, but several things have changed. Libraries now cover the first four layers, so there is a new "build or adopt" section. The code samples are rewritten: the original connection manager retried a POST (which starts a second, billed generation), the error boundary could not catch stream failures, and the token counter counted network chunks instead of tokens. Token counts, cost and context size now come from the server, and the model price table is gone because both models in it are retired. The biggest correction: closing a connection does not stop a generation, so resume and cancel are now separate, server-side operations, and a new section implements that server side.
The Traditional Frontend Playbook Is Broken
You've built dozens of React applications. You know how to manage forms, handle API calls, optimize renders, and structure state. Then you integrate an LLM API and suddenly none of your patterns work. Your component re-renders 50 times per second. Your error boundaries don't catch mid-stream failures. Users navigate away and leave zombie connections consuming tokens. Your carefully optimized bundle now includes streaming parsers, token counters, and connection lifecycle managers you've never needed before.
Figure: Broken Traditional Frontend Playbook
The problem isn't your skills. GenAI fundamentally changed what frontends do. Traditional web apps are request-response systems with discrete state transitions. You submit a form, show a spinner, display the result. State is ephemeral. Errors are atomic. Retries are simple. GenAI apps are real-time streaming systems with continuous state evolution. Tokens arrive at 50Hz. Responses are partial and incremental. Errors happen mid-generation. Connection state matters as much as application state.
This isn't about learning a new library or framework. It's about recognizing that the architectural assumptions underlying your existing patterns no longer hold: synchronous data flow, bounded response times, stateless requests. You're building a different kind of application that requires different architectural thinking.
Mental Model: Frontend as Stream Consumer, Not Request Initiator
The core mental shift is understanding that your frontend is no longer the active party that requests data and waits for a response. It's a passive consumer that subscribes to a data stream, processes events as they arrive, and maintains state across a long-lived connection. This inverts the traditional control flow.
Figure: Frontend as Stream Consumer, Not Request Initiator
In request-response architecture, the frontend controls timing. You send a request when the user clicks submit. You know when the response arrives because the promise resolves. You control when to show loading states, when to update the UI, when to clear state. The lifecycle is deterministic: idle, loading, success, or error.
In streaming architecture, the backend controls timing. It sends tokens whenever the LLM generates them, which might be 10ms apart or 500ms apart depending on model load. Your frontend reacts to events as they arrive. You don't control when tokens come, only how you handle them when they do. The lifecycle is continuous: connecting, streaming, potentially paused by backpressure, eventually completing or erroring.
This mental model affects every architectural decision. State management isn't about "what's the current value" but "what's the accumulated value so far and is the stream still active." Error handling isn't about "did the request fail" but "at what point did the stream fail and what partial data did we successfully receive." Memory management isn't about "clean up after the component unmounts" but "ensure streams are explicitly terminated even when users navigate mid-generation."
The key insight: you're building a real-time system that happens to use HTTP as the transport layer. Treat it like WebSocket communication or video streaming, not like REST API calls. This framing clarifies why your traditional patterns feel inadequate: you're using synchronous tools for asynchronous problems.
Understanding this distinction also reveals why certain patterns are necessary. Connection lifecycle management, explicit stream termination, state accumulation rather than replacement, render batching, and partial response recovery are not optional optimizations. They're fundamental requirements for systems where data arrives continuously and connections persist across time.
One refinement to this model matters more than anything else in the article. The connection and the generation are two different things with two different lifetimes. A generation is a run on your server, calling a model API that bills per token, while a connection is only a view of that run. When the browser closes the connection, the server does not automatically stop the model call. Whether it stops depends on whether your server passes the request's abort signal through to the model call. Vercel's AI SDK documentation says this directly: with stream resumption enabled, closing a tab, refreshing, or calling stop() "closes the current HTTP connection, but it should not cancel the underlying generation."
So design for the server owning the run. In that design, the frontend opens a view of the run, can re-attach to it after a dropped connection, and stops it with an explicit cancel call. Almost every pattern below follows from that one decision.
Architecture: Eight Layers of Frontend Complexity
GenAI frontends require eight distinct architectural layers that traditional applications either don't need or handle trivially. Each layer has its own state, error modes, and performance characteristics.
Architecture: Eight Layers of Frontend Complexity (Open image in new tab, to see the deep details.)
Layer 1: Connection Manager Layer
Handles stream lifecycle from initialization through termination. Tracks connection state, implements reconnection logic, manages timeouts, and ensures cleanup on navigation. This layer is what distinguishes streaming from REST: you need explicit connection state management.
Layer 2: Stream Parser Layer
Transforms raw byte streams into structured events. Handles SSE parsing, manages partial event buffering, deals with protocol-specific quirks like heartbeat events. Without this, you're processing malformed data.
Layer 3: State Accumulator Layer
Aggregates tokens into complete responses. Maintains conversation history, handles multi-turn context, implements optimistic updates. Traditional state management treats updates as replacements. This layer treats them as accumulations.
Layer 4: Render Optimizer Layer
Batches state updates to prevent render storms. Implements debouncing, throttling, or frame-based updates. Without this, 50 token updates per second means 50 component renders per second, killing browser performance.
Layer 5: Token Counter and Cost Tracker Layer
Shows token consumption and cost as the run progresses. Users need live feedback because billing happens per token, not per request. The frontend owns the display, not the numbers. Providers report exact usage in the stream itself: Anthropic's Messages API sends usage on message_start and cumulative counts on message_delta, and OpenAI's Responses API includes usage on the completed response. Your server turns that into tokens and cost, and sends it to the browser as an event. Counting stream chunks in the browser gives you a number that is wrong, and a price table in frontend code goes stale every time a model is retired.
Layer 6: Context Manager Layer
Shows how full the context window is and warns before the limit. Pruning or compacting old messages happens on the server, next to the prompt it changes, and providers increasingly do part of this for you. The frontend displays the utilization the server reports, so users get immediate feedback without the browser needing any model-specific knowledge.
Layer 7: Stream Error Layer
Handles mid-stream failures, rate limit responses, partial recovery scenarios. A React error boundary is the wrong tool for this layer. React's documentation says error boundaries do not catch errors in event handlers or asynchronous code, and a failing fetch loop is asynchronous code. Stream failures have to be modeled as state: which kind of failure, how much partial content arrived, and whether the run is still alive on the server. Keep an error boundary for render errors, such as a markdown renderer crashing on unexpected input.
Layer 8: Memory Manager Layer
Ensures connection cleanup on navigation and prevents memory leaks from abandoned streams. SPAs don't automatically clean up long-lived connections, so you must do it explicitly. This layer also has to keep two operations separate. Disconnect closes the browser's connection and leaves the server-side run alive, so it can be resumed. Cancel tells the server to stop the generation. Unmounting a component should disconnect. Only an explicit user action, or a server-side budget or timeout, should cancel.
Each layer has failure modes that cascade. A missing memory manager causes connection leaks. A naive state accumulator causes render storms. An inadequate error layer leaves users staring at frozen UIs. You can't skip layers or handle them as afterthoughts. They're all load-bearing.
Build or Adopt: Most Teams Should Not Write Layers 1-4
When this article first came out, most teams wrote these layers by hand. That is no longer the sensible default for a chat-shaped interface.
- The Vercel AI SDK's
useChathook (theaipackage, at version 7 as of September 2026) covers connection management, stream parsing and state accumulation, plus render throttling through itsthrottleoption, which is off by default. Withresume: trueand theresumable-streampackage, it re-attaches to a generation that is still running after a page reload. That was listed as future work in the original version of this article. - Streamdown, also from Vercel, is a drop-in replacement for
react-markdownbuilt for streamed output. It handles unterminated code blocks and lists, which is what the hand-written renderer in the original version of this article tried to do. - For agent frontends, where the UI shows tool calls, shared state and approvals rather than only text, CopilotKit and the AG-UI protocol cover layers 1-4 plus agent state. CopilotKit in Production: Where the Abstraction Holds and Where You're on Your Own covers where that abstraction stops, and The Snapshot Tax: Why AG-UI's STATE_DELTA Drifts in Production covers its state-sync failure modes.
Build the layers yourself when your UI is not a chat (a streaming dashboard, a collaborative editor), when you need your own resume and cancel semantics, or when your stream carries events these libraries do not model. Even when you adopt a library, the code below is what it does under the hood, and it shows exactly which guarantees you need to check. Layers 5-8 stay your problem either way, because they depend on your billing, your context policy and your failure modes.
Implementation: Concrete Patterns for Each Layer
Production GenAI frontends need specific implementation patterns for each architectural layer. Here's what actually works.
The examples assume a small server contract. POST /api/runs starts a run and returns a Server-Sent Events (SSE) stream. Every event carries an increasing id. The first event is run with the run ID, text arrives as delta events, the server sends usage events with token counts and cost, refusal or error events report a run that ended badly, and the stream always ends with a done event. GET /api/runs/{id}/stream re-attaches to a live run, starting after the Last-Event-ID header. POST /api/runs/{id}/cancel stops the generation. You will find the server side of this contract implemented at the end of this section.
Connection Manager with Explicit Lifecycle
// hooks/useStreamConnection.tsimport { useCallback, useEffect, useRef, useState } from 'react';export type ConnectionState = 'idle' | 'connecting' | 'streaming' | 'done' | 'error';export interface StreamEvent { id?: string; event: string; data: string;}export class StreamHttpError extends Error { readonly status: number; constructor(status: number) { super(`HTTP ${status}`); this.name = 'StreamHttpError'; this.status = status; }}// Minimal SSE parser. Assumes LF line endings, which common SSE server libraries emit.export async function* parseSSE(body: ReadableStream<Uint8Array>): AsyncGenerator<StreamEvent> { const reader = body.getReader(); const decoder = new TextDecoder(); let buffer = ''; try { while (true) { const { done, value } = await reader.read(); if (done) return; // stream: true keeps a multi-byte character split across chunks intact buffer += decoder.decode(value, { stream: true }); let boundary = buffer.indexOf('\n\n'); while (boundary !== -1) { const block = buffer.slice(0, boundary); buffer = buffer.slice(boundary + 2); boundary = buffer.indexOf('\n\n'); const event: StreamEvent = { event: 'message', data: '' }; const dataLines: string[] = []; for (const line of block.split('\n')) { if (line === '' || line.startsWith(':')) continue; // ':' lines are heartbeats const colon = line.indexOf(':'); const field = colon === -1 ? line : line.slice(0, colon); const value = colon === -1 ? '' : line.slice(colon + 1).replace(/^ /, ''); if (field === 'data') dataLines.push(value); else if (field === 'event') event.event = value; else if (field === 'id') event.id = value; } if (dataLines.length > 0) { event.data = dataLines.join('\n'); yield event; } } } } finally { // Cancel, not just releaseLock: a consumer that stops early must close the connection await reader.cancel().catch(() => {}); }}export function useStreamConnection() { const [state, setState] = useState<ConnectionState>('idle'); const [error, setError] = useState<Error | null>(null); const controllerRef = useRef<AbortController | null>(null); const lastEventIdRef = useRef<string | null>(null); const consume = useCallback( async (open: (signal: AbortSignal) => Promise<Response>, onEvent: (event: StreamEvent) => void) => { controllerRef.current?.abort(); const controller = new AbortController(); controllerRef.current = controller; setState('connecting'); setError(null); try { const response = await open(controller.signal); if (!response.ok || !response.body) throw new StreamHttpError(response.status); setState('streaming'); let sawDone = false; for await (const event of parseSSE(response.body)) { if (event.id) lastEventIdRef.current = event.id; if (event.event === 'done') sawDone = true; onEvent(event); } // A clean end of the body without 'done' means the connection dropped mid-run if (!sawDone) throw new Error('Stream ended before the run finished'); setState('done'); } catch (err) { if (controller.signal.aborted) return; // we closed it on purpose setError(err instanceof Error ? err : new Error(String(err))); setState('error'); } }, [] ); // Starts a new generation. Never retried automatically: a retry is a second, billed run. const start = useCallback( (url: string, payload: unknown, onEvent: (event: StreamEvent) => void) => { lastEventIdRef.current = null; return consume( (signal) => fetch(url, { method: 'POST', signal, headers: { 'Content-Type': 'application/json', Accept: 'text/event-stream' }, body: JSON.stringify(payload), }), onEvent ); }, [consume] ); // Re-attaches to a run that is still going on the server. Safe to retry: no new generation. const resume = useCallback( (streamUrl: string, onEvent: (event: StreamEvent) => void) => { const headers: Record<string, string> = { Accept: 'text/event-stream' }; if (lastEventIdRef.current) headers['Last-Event-ID'] = lastEventIdRef.current; return consume((signal) => fetch(streamUrl, { signal, headers }), onEvent); }, [consume] ); // Closes the connection only. The server-side run keeps going and can be resumed. const disconnect = useCallback(() => { controllerRef.current?.abort(); controllerRef.current = null; setState('idle'); }, []); // Stops the generation itself. Closing the connection is not enough to stop billing. const cancel = useCallback( async (cancelUrl: string) => { disconnect(); await fetch(cancelUrl, { method: 'POST' }); }, [disconnect] ); useEffect(() => () => controllerRef.current?.abort(), []); return { state, error, start, resume, disconnect, cancel };}
Three decisions in this hook fix bugs that are common in hand-written streaming code, including the first version of this article.
A POST is never retried automatically. Retrying the request that starts a generation starts a second generation. The user pays twice, and a retry after 300 tokens restarts the answer from the beginning. start runs once. resume is the retry path: it is a GET against a run that already exists, it sends Last-Event-ID, and the server continues from the next event.
A connection that ends without done is an error. A mobile network switch or a proxy timeout often ends the response body cleanly, with no exception. Code that treats the end of the body as success shows a truncated answer as if it were complete. Requiring an explicit terminal event turns that silent truncation into a Connection interrupted state the user can act on.
The decoder runs with stream: true, and events are parsed, not forwarded. A network chunk can end in the middle of a multi-byte UTF-8 character, such as an emoji or most non-Latin text, or in the middle of an SSE event. Decoding each chunk on its own produces replacement characters, and forwarding raw chunks passes half-events to the rest of the app. Instead, the parser buffers until a blank line ends an event. It was tested with a stream delivered one byte at a time. When the consumer stops reading early, the finally block cancels the reader instead of only releasing its lock, so the connection actually closes and the server stops sending.
disconnect and cancel are separate on purpose. Unmounting disconnects, and the run stays alive for resume. The Stop button cancels, which is the only thing that stops billing. On the server, the model call must not be tied to the request's abort signal, or every dropped connection kills the run you are trying to resume.
State Accumulator with Render Batching
// hooks/useStreamingState.tsimport { useState, useRef, useCallback, useEffect } from 'react';export function useStreamingState() { const [displayContent, setDisplayContent] = useState(''); const bufferRef = useRef(''); const rafRef = useRef<number | null>(null); const appendToken = useCallback((token: string) => { bufferRef.current += token; // Batch updates using requestAnimationFrame if (!rafRef.current) { rafRef.current = requestAnimationFrame(() => { setDisplayContent(bufferRef.current); rafRef.current = null; }); } }, []); const flush = useCallback(() => { if (rafRef.current) { cancelAnimationFrame(rafRef.current); rafRef.current = null; } setDisplayContent(bufferRef.current); }, []); const reset = useCallback(() => { bufferRef.current = ''; setDisplayContent(''); if (rafRef.current) { cancelAnimationFrame(rafRef.current); rafRef.current = null; } }, []); // Cleanup on unmount useEffect(() => { return () => { if (rafRef.current) { cancelAnimationFrame(rafRef.current); } }; }, []); return { content: displayContent, appendToken, flush, reset };}
This batches rapid token updates to match the display refresh rate (typically 60Hz). Without batching, 50 token updates per second causes 50 React renders, 50 virtual DOM diffs, and 50 DOM updates. None of that is needed, because humans can't perceive updates faster than 16ms anyway.
Server-Reported Usage and Context
The original version of this article counted tokens in the browser, with a hardcoded price table. It had three problems. It counted stream chunks, and a chunk is not a token. It priced output tokens only. And both models in its price table have since been retired (claude-3-opus-20240229 on January 5, 2026, and gpt-4 is scheduled for shutdown on October 23, 2026). The same applies to the context window: an 8,000-token default is far below current models, and pruning messages in the browser changes a prompt the browser does not own.
Move the numbers to the server, which already has them. The server reads the provider's reported usage, applies the current price for the model it actually called, and knows how large the context it sent was. It sends the result as a usage event:
id: 42event: usagedata: {"inputTokens":2679,"outputTokens":510,"costUsd":0.0412,"contextTokens":3189,"contextLimit":200000}
The frontend only stores and displays it:
// hooks/useRunMetrics.tsimport { useCallback, useState } from 'react';// Computed on the server from the provider's reported usage, never estimated in the browser.export interface RunMetrics { inputTokens: number; outputTokens: number; costUsd: number; contextTokens: number; contextLimit: number;}export function useRunMetrics(warnAt = 0.8) { const [metrics, setMetrics] = useState<RunMetrics | null>(null); const applyUsageEvent = useCallback((data: string) => { setMetrics(JSON.parse(data) as RunMetrics); }, []); const utilization = metrics ? metrics.contextTokens / metrics.contextLimit : 0; return { metrics, applyUsageEvent, utilization, isNearContextLimit: utilization >= warnAt, };}
This is the architecture a reader proposed in the comments below: the backend computes token counts, context utilization and cost, and the frontend subscribes and renders. It is now the default recommendation, not a multi-team special case. When isNearContextLimit is true, the UI can warn and offer to start a new conversation or summarize the old one. Any summarizing itself happens on the server.
Progressive Markdown Rendering
Streamed markdown breaks in predictable ways: a code fence that has opened but not closed, a list with half an item, a link without its closing parenthesis. The original version of this article included a hand-written renderer that held back incomplete blocks. It re-parsed the entire response on every animation frame, so the work grew with the square of the response length.
Use Streamdown instead. It is a drop-in replacement for react-markdown built for streamed content. It repairs unterminated blocks as they arrive, memoizes rendering so completed content is not re-rendered on every update, and hardens the output HTML, which matters because model output is untrusted input.
import { Streamdown } from 'streamdown';<Streamdown isAnimating={state === 'streaming'}>{content}</Streamdown>
Streamdown's styles rely on Tailwind scanning its package files and on shadcn/ui-style CSS variables. Check its installation notes before assuming it will look right in an existing design system.
Stream Errors as State, Not Error Boundaries
The original version of this article classified stream failures in a React error boundary. That component could never have run for a stream failure. React's documentation lists asynchronous code among the things error boundaries do not catch, and the fetch loop is asynchronous. The failure has to be read from the connection hook's state instead:
// components/StreamStatus.tsx'use client';import { StreamHttpError, type ConnectionState } from '../hooks/useStreamConnection';interface StreamStatusProps { state: ConnectionState; error: Error | null; hasPartialContent: boolean; onResume: () => void; onRegenerate: () => void;}function describe(error: Error | null) { if (error instanceof StreamHttpError) { if (error.status === 429) { return { title: 'Rate limit reached', detail: 'Wait a moment before regenerating.', canResume: false }; } if (error.status >= 500) { return { title: 'The model service failed', detail: 'This is on our side, not your network.', canResume: false }; } return { title: 'Request rejected', detail: `The server returned ${error.status}.`, canResume: false }; } // No HTTP status: the connection dropped after the run started, so the run may still be alive return { title: 'Connection interrupted', detail: 'The response may still be generating.', canResume: true };}export function StreamStatus({ state, error, hasPartialContent, onResume, onRegenerate }: StreamStatusProps) { if (state !== 'error') return null; const { title, detail, canResume } = describe(error); return ( <div role="alert" className="p-4 bg-red-50 border border-red-200 rounded"> <p className="font-semibold text-red-800">{title}</p> <p className="text-red-600 mt-2"> {detail} {hasPartialContent ? ' The partial response above has been kept.' : ''} </p> <div className="mt-4 flex gap-2"> {canResume && ( <button onClick={onResume} className="px-4 py-2 bg-red-600 text-white rounded"> Resume </button> )} <button onClick={onRegenerate} className="px-4 py-2 border border-red-600 text-red-700 rounded"> Regenerate (new run) </button> </div> </div> );}
The classification follows what the user can actually do next. A 429 or a 5xx happened before the run produced anything useful, so the only option is a new run, labelled as a new run because it is billed again. An error with no HTTP status means the connection dropped after the run started, so the run may still be alive and Resume is the first choice. In both cases the partial content stays on screen.
Wiring It Together
// components/ChatResponse.tsx'use client';import { useCallback, useRef, useState } from 'react';import { Streamdown } from 'streamdown';import { useStreamConnection, type StreamEvent } from '../hooks/useStreamConnection';import { useStreamingState } from '../hooks/useStreamingState';import { useRunMetrics } from '../hooks/useRunMetrics';import { StreamStatus } from './StreamStatus';export function ChatResponse({ prompt }: { prompt: string }) { const { state, error, start, resume, disconnect, cancel } = useStreamConnection(); const { content, appendToken, flush, reset } = useStreamingState(); const { metrics, applyUsageEvent, isNearContextLimit } = useRunMetrics(); const runIdRef = useRef<string | null>(null); const [notice, setNotice] = useState<string | null>(null); const onEvent = useCallback( (event: StreamEvent) => { switch (event.event) { case 'run': runIdRef.current = (JSON.parse(event.data) as { runId: string }).runId; break; case 'delta': appendToken((JSON.parse(event.data) as { text: string }).text); break; case 'usage': applyUsageEvent(event.data); break; case 'refusal': setNotice('The model declined this request.'); break; case 'error': { const { reason } = JSON.parse(event.data) as { reason: string }; if (reason !== 'cancelled') setNotice('Generation failed on the server. The partial response has been kept.'); break; } case 'done': flush(); break; } }, [appendToken, applyUsageEvent, flush] ); const send = () => { reset(); setNotice(null); runIdRef.current = null; void start('/api/runs', { prompt }, onEvent); }; // Resumed events continue from Last-Event-ID, so the buffer is appended to, not reset const reattach = () => { if (runIdRef.current) void resume(`/api/runs/${runIdRef.current}/stream`, onEvent); }; const stop = () => { flush(); if (runIdRef.current) void cancel(`/api/runs/${runIdRef.current}/cancel`); else disconnect(); // no run ID yet: nothing on the server to cancel by ID }; const isActive = state === 'connecting' || state === 'streaming'; return ( <div> <Streamdown isAnimating={state === 'streaming'}>{content}</Streamdown> {notice && <p role="status">{notice}</p>} <StreamStatus state={state} error={error} hasPartialContent={content.length > 0} onResume={reattach} onRegenerate={send} /> {metrics && ( <p className={isNearContextLimit ? 'text-amber-700' : 'text-gray-500'}> {metrics.inputTokens + metrics.outputTokens} tokens, context {metrics.contextTokens} / {metrics.contextLimit} </p> )} <button onClick={isActive ? stop : send}>{isActive ? 'Stop generating' : 'Send'}</button> </div> );}
Note what Stop generating does: it flushes what has already arrived and calls the server's cancel endpoint. It does not only abort the fetch. A Stop button that only aborts the fetch looks like it works, because the text stops appearing, while the server-side run keeps generating tokens you pay for.
The Server Side: Runs That Outlive Connections
Everything above depends on the server holding up its side of the contract. What matters most is that a run is not a request. It has its own lifetime, its own abort controller and its own event log, and connections come and go while it runs.
// lib/runs.ts// In-memory run registry for a single server process. With more than one instance,// keep the event log in Redis instead (the AI SDK's resumable-stream package does this).export type Emit = (event: string, data: unknown) => void;type Producer = (emit: Emit, signal: AbortSignal) => Promise<void>;interface RunEvent { id: number; event: string; data: string;}export interface Run { id: string; ownerId: string; events: RunEvent[]; finished: boolean; controller: AbortController; listeners: Set<() => void>;}export const SSE_HEADERS = { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache, no-transform', 'X-Accel-Buffering': 'no', // stops nginx from buffering the stream};const RUN_TIMEOUT_MS = 5 * 60_000;const KEEP_FINISHED_RUN_MS = 10 * 60_000;const HEARTBEAT_MS = 15_000;const runs = new Map<string, Run>();export function startRun(ownerId: string, produce: Producer): Run { const run: Run = { id: crypto.randomUUID(), ownerId, events: [], finished: false, controller: new AbortController(), listeners: new Set(), }; runs.set(run.id, run); const emit: Emit = (event, data) => { run.events.push({ id: run.events.length + 1, event, data: JSON.stringify(data) }); run.listeners.forEach((notify) => notify()); }; // The run owns its own AbortController, not the request's signal, // so it outlives the connection that started it. The timeout is its budget. const timeout = setTimeout(() => run.controller.abort(), RUN_TIMEOUT_MS); emit('run', { runId: run.id }); produce(emit, run.controller.signal) .catch((err: unknown) => { if (!run.controller.signal.aborted) console.error(`run ${run.id} failed`, err); emit('error', { reason: run.controller.signal.aborted ? 'cancelled' : 'generation_failed' }); }) .finally(() => { clearTimeout(timeout); run.finished = true; emit('done', {}); setTimeout(() => runs.delete(run.id), KEEP_FINISHED_RUN_MS); }); return run;}export function getRun(id: string, ownerId: string | null): Run | undefined { const run = runs.get(id); return run && run.ownerId === ownerId ? run : undefined;}export function cancelRun(run: Run): void { run.controller.abort();}// Streams every event after `afterId`, then follows the run live until it finishes.export function subscribe(run: Run, afterId: number): ReadableStream<Uint8Array> { const encoder = new TextEncoder(); let cursor = afterId; let notify: (() => void) | undefined; let heartbeat: ReturnType<typeof setInterval> | undefined; const stopFollowing = () => { if (notify) run.listeners.delete(notify); clearInterval(heartbeat); }; return new ReadableStream<Uint8Array>({ start(controller) { notify = () => { for (const e of run.events.slice(cursor)) { controller.enqueue(encoder.encode(`id: ${e.id}\nevent: ${e.event}\ndata: ${e.data}\n\n`)); cursor = e.id; } if (run.finished && cursor === run.events.length) { stopFollowing(); controller.close(); } }; run.listeners.add(notify); heartbeat = setInterval(() => controller.enqueue(encoder.encode(': ping\n\n')), HEARTBEAT_MS); notify(); }, // The client disconnected. Stop sending to it, and leave the run alone. cancel() { stopFollowing(); }, });}
The registry does four things that a plain streaming route handler does not.
It owns the abort controller. The model call is tied to run.controller, never to request.signal. A dropped connection ends one subscription and nothing else. The run ends when the model finishes, when the user cancels, or when the 5-minute timeout fires. That timeout is the budget that stops an abandoned run from generating forever.
It keeps an event log with increasing IDs. subscribe replays every event after afterId, then follows the run live. That one function serves both the first connection (afterId of 0) and every resume (the client's Last-Event-ID), so a resumed client gets exactly the events it missed, with no duplicates.
It checks ownership on every lookup. A run ID alone must not be enough to read someone's conversation or cancel their run, so getRun returns nothing for another user's run, and the routes answer 404 rather than 403, so they don't confirm the run exists.
It keeps the stream alive through proxies. A comment line every 15 seconds stops idle-connection timeouts during a long pause, such as a model thinking before its first token. The X-Accel-Buffering: no header stops nginx from buffering the whole response, which would turn a stream into one late delivery.
The model call reads usage from the stream and turns it into the usage events the frontend displays:
// lib/generate.tsimport Anthropic from '@anthropic-ai/sdk';import type { Emit } from './runs';const client = new Anthropic();// Prices per million tokens and context sizes live on the server, so a pricing change// or a model retirement is one server deploy, not a frontend release.// Cache writes are priced at the 5-minute rate for brevity.const MODELS: Record<string, { input: number; cacheWrite: number; cacheRead: number; output: number; contextLimit: number }> = { 'claude-opus-5': { input: 5, cacheWrite: 6.25, cacheRead: 0.5, output: 25, contextLimit: 1_000_000 }, 'claude-opus-4-8': { input: 5, cacheWrite: 6.25, cacheRead: 0.5, output: 25, contextLimit: 1_000_000 },};interface Usage { input_tokens: number; cache_creation_input_tokens: number; cache_read_input_tokens: number; output_tokens: number;}function toMetrics(model: string, usage: Usage) { const p = MODELS[model] ?? MODELS['claude-opus-5']; const costUsd = (usage.input_tokens * p.input + usage.cache_creation_input_tokens * p.cacheWrite + usage.cache_read_input_tokens * p.cacheRead + usage.output_tokens * p.output) / 1_000_000; const inputTokens = usage.input_tokens + usage.cache_creation_input_tokens + usage.cache_read_input_tokens; return { inputTokens, outputTokens: usage.output_tokens, costUsd, contextTokens: inputTokens + usage.output_tokens, contextLimit: p.contextLimit, };}export async function generate(prompt: string, emit: Emit, signal: AbortSignal): Promise<void> { const stream = client.beta.messages.stream( { model: 'claude-opus-5', max_tokens: 64000, // If the model declines, the API re-runs the request on a fallback model in the same stream betas: ['server-side-fallback-2026-07-01'], fallbacks: 'default', messages: [{ role: 'user', content: prompt }], }, { signal } ); let model = 'claude-opus-5'; let startUsage: Usage = { input_tokens: 0, cache_creation_input_tokens: 0, cache_read_input_tokens: 0, output_tokens: 0 }; for await (const event of stream) { switch (event.type) { case 'message_start': { model = event.message.model; const u = event.message.usage; startUsage = { input_tokens: u.input_tokens, cache_creation_input_tokens: u.cache_creation_input_tokens ?? 0, cache_read_input_tokens: u.cache_read_input_tokens ?? 0, output_tokens: u.output_tokens, }; emit('usage', toMetrics(model, startUsage)); break; } case 'content_block_delta': if (event.delta.type === 'text_delta') emit('delta', { text: event.delta.text }); break; case 'message_delta': // Token counts on message_delta are cumulative emit('usage', toMetrics(model, { ...startUsage, output_tokens: event.usage.output_tokens })); break; } } const final = await stream.finalMessage(); emit( 'usage', toMetrics(final.model, { input_tokens: final.usage.input_tokens, cache_creation_input_tokens: final.usage.cache_creation_input_tokens ?? 0, cache_read_input_tokens: final.usage.cache_read_input_tokens ?? 0, output_tokens: final.usage.output_tokens, }) ); // A refusal here means the model and its fallback both declined if (final.stop_reason === 'refusal') emit('refusal', { category: final.stop_details?.category ?? null });}
This uses Claude Opus 5 through Anthropic's TypeScript SDK. Usage comes from the provider, not from counting: message_start carries input tokens, and message_delta carries cumulative output tokens. At the end, finalMessage() provides the authoritative totals. Prices and context sizes sit in a server-side table (checked against Anthropic's pricing page in September 2026), so a price change is a server deploy. The request also enables server-side fallbacks: if the model's safety classifiers decline a request, the API re-runs it on a fallback model inside the same stream, and message_start names whichever model actually served it. A refusal stop reason at the end means both declined, and the frontend shows that instead of an empty answer.
Each route is thin. Starting a run:
// app/api/runs/route.tsimport { getUserId } from '@/lib/auth'; // your existing session lookupimport { generate } from '@/lib/generate';import { SSE_HEADERS, startRun, subscribe } from '@/lib/runs';export const runtime = 'nodejs'; // the run registry lives in this process's memoryexport async function POST(request: Request) { const userId = await getUserId(request); if (!userId) return new Response(null, { status: 401 }); const { prompt } = (await request.json()) as { prompt?: unknown }; if (typeof prompt !== 'string' || prompt.length === 0) return new Response(null, { status: 400 }); // request.signal is deliberately not passed to the run: a dropped connection must not end it const run = startRun(userId, (emit, signal) => generate(prompt, emit, signal)); return new Response(subscribe(run, 0), { headers: SSE_HEADERS });}
Re-attaching to a run:
// app/api/runs/[id]/stream/route.tsimport { getUserId } from '@/lib/auth';import { SSE_HEADERS, getRun, subscribe } from '@/lib/runs';export const runtime = 'nodejs';export async function GET(request: Request, { params }: { params: Promise<{ id: string }> }) { const { id } = await params; const run = getRun(id, await getUserId(request)); if (!run) return new Response(null, { status: 404 }); const afterId = Number(request.headers.get('Last-Event-ID') ?? 0) || 0; return new Response(subscribe(run, afterId), { headers: SSE_HEADERS });}
Cancelling one:
// app/api/runs/[id]/cancel/route.tsimport { getUserId } from '@/lib/auth';import { cancelRun, getRun } from '@/lib/runs';export const runtime = 'nodejs';export async function POST(request: Request, { params }: { params: Promise<{ id: string }> }) { const { id } = await params; const run = getRun(id, await getUserId(request)); if (!run) return new Response(null, { status: 404 }); cancelRun(run); // aborts the model call, which is what stops the billing return new Response(null, { status: 204 });}
These use the Next.js App Router, where params is a Promise. The registry is in-process memory, so it works on one long-running Node.js server. Serverless functions and multiple instances do not share memory. There, move the event log and the cancel flag to Redis, which is what the AI SDK's resumable-stream package does. Nothing changes in the contract the frontend sees.
This whole path was tested end to end with a fake model generator in place of the API call. A client disconnected after two deltas while the run kept producing. It resumed with Last-Event-ID and reassembled the full text with no duplicate event IDs. Cancel stopped the generator early and reported cancelled, and a subscriber that connected after the run finished got a full replay.
Pitfalls and Failure Modes
Render Storms from Naive State Updates
The most common mistake is calling setState on every token arrival. At 50 tokens per second, this means 50 React renders per second. The browser can't keep up. UI becomes laggy. Memory usage spikes. Eventually, the tab freezes.
Detection: open DevTools performance profiler during streaming. If you see continuous render cycles consuming >80% of frame time, you have a render storm. Solution: batch updates with requestAnimationFrame or debounce with fixed intervals (50-100ms).
Memory Leaks from Unclosed Streams
Users navigate away mid-generation. The component unmounts but the fetch request continues. The stream stays open, consuming memory. In SPAs where users navigate frequently, this accumulates dozens of zombie connections.
Detection: monitor network tab for streams that continue after navigation. Check browser memory profiler for growing heap even when idle. Solution: always implement cleanup in useEffect return functions. Use AbortController to close fetch requests on unmount.
Closing the connection fixes the browser's memory. It does not necessarily stop the model call on your server, which is where the tokens are spent. That is the next pitfall.
Treating Disconnect as Cancel
The fetch is aborted, the text stops appearing, and the Stop button looks like it works. On the server, the model call keeps running to its token limit, because nothing told it to stop. Or the reverse: the server does tie the model call to the request's abort signal, and every mobile network switch kills a run the user expected to come back to.
Detection: stop a long generation from the UI, then check provider usage or server logs for that run's final output token count. If it keeps growing after Stop, disconnect is not cancelling. Solution: decide per action. Unmount and network loss disconnect, and the run stays resumable. Stop calls a server cancel endpoint. Put a server-side token budget and timeout on every run, so an abandoned run still ends.
Retrying the Request That Starts a Generation
A stream fails after 300 tokens, and generic retry logic re-sends the POST. The model starts a second generation from the beginning. The user sees the answer restart, and you pay for both runs. With exponential backoff and three attempts, one flaky connection can cost four generations.
Detection: count runs per user message in your server logs. Anything above one without an explicit user action is an automatic retry. Solution: never retry the POST automatically. Resume the existing run with a GET and Last-Event-ID. Offer a new generation only as an explicit, labelled user action.
Context Window Overflow Without Warning
Users have multi-turn conversations. Each turn adds tokens. Eventually they hit the context limit. The LLM request fails with a cryptic error. Users don't understand why. They just see "request failed."
Detection: have the server report context usage for every run. Solution: show a context utilization meter in the UI, driven by the server's numbers. Warn at 80% capacity. Summarize or prune old messages on the server, or offer the user a fresh conversation, before the limit is reached.
Partial Response Loss on Error
Stream fails after receiving 500 tokens. Traditional error handling clears state and shows an error message. The user loses the partial response that might have been useful.
Detection: check if error handlers wipe state unconditionally. Solution: preserve partial content on stream errors. Show error message alongside partial response. Give users option to retry or accept partial result.
Cost Explosion from Background Streaming
User opens multiple tabs or leaves tabs open. Each tab maintains its own stream. Forgotten tabs continue consuming tokens. User racks up hundreds of dollars before noticing.
Detection: implement cost tracking per session on the server. Alert when cost exceeds thresholds. Solution: show a persistent cost indicator in the UI. Enforce budgets on the server, not in the tab: a per-run token limit, a per-session cost limit, and a timeout for runs nobody is watching. The original version of this article suggested pausing streams with the Page Visibility API when a tab is hidden. That pauses the display, not the generation. An LLM stream cannot be paused on the provider side, so the tokens are generated and billed whether the tab is visible or not.
Error Boundaries That Never Fire
The team wraps the chat in an error boundary and considers stream failures handled. Then a connection drops mid-response and the UI freezes on half an answer, with no error shown.
Detection: kill the network in DevTools during a stream. If no error UI appears, the boundary is not catching it. Solution: error boundaries only catch errors thrown during rendering, not in asynchronous code like a fetch loop. Model stream failures as state in the connection hook, as shown above, and keep the boundary for render crashes.
Broken Markdown Rendering During Streaming
Markdown parser tries to render incomplete code blocks. UI flickers as it alternates between code block formatting and plain text. Lists break when only partial items have arrived.
Detection: watch for UI flickering during streaming. Check if markdown elements appear and disappear. Solution: use a renderer built for streamed markdown, such as Streamdown, which repairs unterminated blocks as they arrive instead of re-parsing the whole response on every update.
Summary and Next Steps
GenAI frontends require fundamentally different architectural patterns than traditional web applications. The shift from request-response to streaming inverts control flow, makes connection state explicit, and requires eight distinct architectural layers: connection management, stream parsing, state accumulation, render optimization, token tracking, context management, error boundaries, and memory management.
The key insights: treat streaming as a protocol, not an HTTP response variant, and implement explicit lifecycle management. Separate the connection from the generation: the server owns the run, the browser disconnects and resumes, and only an explicit action cancels. Never retry the request that starts a generation. Batch renders to prevent storms: 50 token updates per second needs batching, not 50 React renders. Show tokens, cost and context utilization live, but compute them on the server from provider-reported usage. Model stream errors as state, because error boundaries do not see them. Preserve partial responses on errors, because 500 tokens of useful output is better than nothing. Clean up connections on navigation, because SPAs don't automatically terminate streams.
Next steps for production: for chat-shaped interfaces, start from the Vercel AI SDK and Streamdown before writing layers 1-4 yourself, and use the code above to check what they guarantee. Implement comprehensive observability for stream health: track connection duration, token throughput, error rates, resume rates, and runs per user message. Put token budgets and timeouts on every server-side run. Build A/B testing infrastructure to measure streaming UX impact. Does it improve engagement, or does it only feel faster? Implement graceful degradation by polling the run's status when SSE fails. Build cost prediction models that warn users before expensive operations.
The patterns described here work for current LLM APIs but will need adaptation as models get faster, responses get longer, and multi-modal streaming becomes common. Agent frontends already stream more than text: tool calls, structured output and shared state, which is where protocols like AG-UI come in. Stay focused on fundamentals: explicit state management, connection lifecycle control, progressive rendering, and user feedback. These principles survive API changes.
References
- Vercel. AI SDK UI: Chatbot Resume Streams. https://ai-sdk.dev/docs/ai-sdk-ui/chatbot-resume-streams
- Vercel. Troubleshooting: Abort and resumable streams. https://ai-sdk.dev/docs/troubleshooting/abort-breaks-resumable-streams
- Vercel. Streamdown. https://github.com/vercel/streamdown
- React. Component: catching rendering errors with an error boundary. https://react.dev/reference/react/Component
- Anthropic. Streaming messages. https://platform.claude.com/docs/en/build-with-claude/streaming
- Anthropic. Model deprecations. https://platform.claude.com/docs/en/about-claude/model-deprecations
- Anthropic. Refusals and fallback. https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback
- Anthropic. Pricing. https://platform.claude.com/docs/en/about-claude/pricing
- Anthropic. Models overview. https://platform.claude.com/docs/en/about-claude/models/overview
- Anthropic. TypeScript SDK. https://github.com/anthropics/anthropic-sdk-typescript
- Next.js. Route Handlers. https://nextjs.org/docs/app/api-reference/file-conventions/route
- OpenAI. Streaming API responses. https://developers.openai.com/api/docs/guides/streaming-responses
- OpenAI. Deprecations. https://developers.openai.com/api/docs/deprecations
- MDN. TextDecoder: decode() method. https://developer.mozilla.org/en-US/docs/Web/API/TextDecoder/decode
- WHATWG. HTML Living Standard: Server-sent events. https://html.spec.whatwg.org/multipage/server-sent-events.html
Comments
Following is the valuable comments/questions from Yuvraj Shivaji Dhepe
Question 1:
For the first 4 layers, will using CopilotKit be helpful, it's something like frontend for such agentic systems, they have major integrations with agentic frameworks like Agno, LangGraph etc.
Answer
CopilotKit for Layers 1-4
Answer is Yes and No (It depends)
Yes, CopilotKit abstracts away significant complexity such as connection management, stream parsing, state accumulation, some render optimization and provides React hooks..
When it's worth it:
- You're building standard chat interfaces
- You want rapid prototyping
- Your frontend team isn't deep in streaming internals
- You're okay with opinionated patterns (Meaning that you don't want significant deviations from what is already there in terms of user user experience/journey).
When you'll outgrow it:
- Custom streaming protocols (not standard SSE/WebSocket)
- Fine-grained control over reconnection logic
- Non-chat UIs (streaming dashboards, collaborative editors)
- Performance optimization beyond what CopilotKit provides
- Multi-provider support with custom handling
- Enterprise environment where genAI is just one part of the overall systems
CopilotKit is essentially "batteries included" for the first 4 layers. If it fits your use case, use it. If you need custom behavior, you'll implement those layers yourself anyway.
Update (September 2026): I went deeper on where these limits fall in CopilotKit in Production: Where the Abstraction Holds and Where You're on Your Own.
Question 2:
For layer 5, 6: I was wondering handling these as a common state between frontend and backend. So a websocket connection where backend just updates: context length left/filled, token counter. Will this not be optimal, in cases when there are multiple teams involved: For example in my case our team is involved in agentic part and frontend team would receive all the LLM related information from the backend itself.
Answer
Backend-as-State-Source via WebSocket (Layers 5-6)
Yes, this is not just optimal, it's the correct architecture for multi-team setups.
Your approach as I understood can be described by the following diagram:
Figure: Backend-as-State-Source via WebSocket
Backend computes:
- Token counts (actual, not estimated)
- Context utilization
- Cost accumulation
- Remaining capacity
Middleware maintains:
- Current state snapshot
- Historical metrics
- Per-session tracking
Frontend subscribes:
- Reactive updates when state changes
- No LLM logic required
- Just display what middleware provides
Why this is optimal:
-
Clean separation of concerns
- Agentic team: owns token logic, context management, LLM interactions
- Frontend team: displays UI, no domain knowledge needed
- Middleware: state synchronization layer
-
Single source of truth
- Backend has actual token counts from provider APIs
- Frontend can't drift out of sync with reality
- No "frontend estimates vs backend actual" mismatches
-
Real-time synchronization
- WebSocket pushes updates immediately
- Frontend always shows current state
- No polling, no stale data
-
Scalability
- Multiple frontend clients can subscribe to same state
- Middleware can broadcast to all connected clients
- Backend doesn't care how many frontends exist
Implementation pattern:
// Backend sends via WebSocket{ type: "metrics_update", session_id: "sess_123", data: { tokens: { input: 245, output: 1247, total: 1492, limit: 8000, remaining: 6508 }, context: { messages: 12, tokens_used: 1492, utilization_percent: 18.7, can_add_message: true }, cost: { current_session: 0.045, estimated_next_message: 0.003 } }}// Frontend just consumesconst [metrics, setMetrics] = useState(null);useEffect(() => { ws.on('metrics_update', (data) => { setMetrics(data); });}, []);// Display<div> <p>Tokens: {metrics.tokens.total} / {metrics.tokens.limit}</p> <ProgressBar value={metrics.context.utilization_percent} /> <p>Cost: \${metrics.cost.current_session}</p></div>
This eliminates frontend complexity:
- No token counting logic
- No context window calculations
- No cost estimation formulas
- No model-specific knowledge
- Just subscribe and render
For your multi-team scenario, this is ideal:
- Agentic team controls the state computation
- Frontend team has simple reactive UI
- Clear contract: WebSocket message schema
- Teams can work independently
The middleware/state manager can be:
- Part of your backend (simple approach)
- Separate service (if you need caching, replay, etc.)
- Redis/similar with pub/sub (for multi-instance backends)
Your instinct is correct - backend-managed state via WebSocket is the right pattern when teams are separated and frontend shouldn't have LLM domain logic.
Question 3:
For layer 7, 8: Again I think this heavy lifting is done via CopilotKit, but haven't looked under the hood?
Answer
CopilotKit for Layers 7-8
Layer 7 (Error Boundaries): CopilotKit handles basic error boundaries but not production-grade classification. It catches stream failures but doesn't differentiate:
- Rate limits (needs backoff UI)
- Context overflow (needs pruning suggestion)
- Network interruption (needs retry)
- Provider outage (needs fallback)
You'll still need custom error boundaries for sophisticated error handling.
Layer 8 (Memory Management): CopilotKit does handle cleanup on unmount. This part works well. Their hooks properly abort connections and clean up listeners.
What CopilotKit doesn't handle:
- Cross-tab coordination (multiple tabs streaming)
- Partial response recovery on error
- Custom reconnection strategies
- Stream prioritization under load
For comments/feedback on the article and to connect with me for more insights on Agentic AI systems, security, and practical AI engineering:
Related Articles
More Articles
- Context Engineering: The Skill That Separates Production Agents from Demos
- State Management for Agentic Systems: How to Build Agents That Don't Start Over
- LangGraph Checkpoints Restore Your Limits, Not Just Your State
Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:




