AI Coding
Stream Model Output from a Next.js Route Handler
Proxy a streaming model response through a Next.js route so the browser can render tokens as they arrive.
- Next.js
- Streaming
- Route Handlers
- AI
On this page
A model that takes twenty seconds to answer feels frozen if your server waits for the last token before it sends the first byte. Streaming means you forward chunks as the provider sends them. The browser can paint the paragraph while the rest is still generating. In the App Router, a route handler can return a Response whose body is a stream. You do not have to buffer the provider's reply into a string first.
Keep the provider key on the server. The browser calls your route. Your route checks the user, checks the size of the prompt, calls the provider with stream set to true, and returns the upstream body. If you buffer with await upstream.text(), you have built a slow non-streaming proxy and paid the latency you meant to avoid.
Forward the body
The exact JSON differs by provider. The shape of the route does not. You POST to your handler, the handler POSTs to the provider, and the handler returns upstream.body with the content type the client expects. For many chat APIs that is text/event-stream. Do not set a content length. You do not know it yet.
export const maxDuration = 60;
export async function POST(request: Request) {
const { prompt } = await request.json();
if (typeof prompt !== "string" || prompt.length < 1 || prompt.length > 8000) {
return Response.json({ error: "Prompt length is out of range." }, { status: 400 });
}
const upstream = await fetch("https://api.openai.com/v1/responses", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.OPENAI_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: process.env.OPENAI_MODEL,
input: prompt,
stream: true,
}),
});
if (!upstream.ok || !upstream.body) {
return Response.json({ error: "Model request failed." }, { status: 502 });
}
return new Response(upstream.body, {
headers: {
"Content-Type": "text/event-stream; charset=utf-8",
"Cache-Control": "no-cache, no-transform",
},
});
}
Confirm the provider path and the body fields in the current docs. The Responses API and the older chat completions API do not take the same JSON. The model id comes from the environment so you can change it without editing the route. If OPENAI_API_KEY is missing, fail the request. A fallback to a fake completion will be mistaken for a working product.
Read the stream in the browser
fetch returns a body you can read with getReader. Decode chunks as UTF-8 and append them. If the provider speaks server-sent events, split on blank lines and parse the data field. Show a stop button that aborts an AbortController. A stream with no cancel is a tab that keeps spending quota after the user has left.
export async function readTextStream(response: Response, onChunk: (text: string) => void) {
const reader = response.body?.getReader();
if (!reader) throw new Error("No stream");
const decoder = new TextDecoder();
while (true) {
const { value, done } = await reader.read();
if (done) break;
onChunk(decoder.decode(value, { stream: true }));
}
}
Guard the route
- Require a session. An open route is an open bill.
- Limit prompt length and request rate. Streaming does not make abuse cheaper.
- Do not log the full prompt if it can contain private source code from a customer.
- Map provider errors to a short message. Forwarding the raw body can leak account details.
- Set maxDuration on platforms that kill long functions, and still tell the user when you hit it.
Caching a streamed completion is usually wrong. The answer depends on the prompt, and a cached stream can be half an event. Send no-store or no-cache. If you want to save the final text, write it after the stream ends, on the server, under the user's id. Do not cache it on a shared CDN URL.
Failure looks like silence
The ugly failures are a 200 status with an error event inside the stream, a proxy that buffers despite your headers, and a client that waits for the stream to close before it renders. Test with a short prompt and watch the first paint. Then test with the provider key removed and confirm the user sees a 502, not an infinite spinner. Streaming is a transport. The product is still a request that can fail.
Proxies will try to buffer you
A platform, a compression middleware, or a well-meaning helper can collect the stream and send it at the end. If the first token arrives only when the model finishes, something in the path is buffering. Disable compression for that route if your host applies it by default. Avoid wrapping upstream.body in a function that reads it to a string for logging. Log the status and the duration, not the tokens. On the client, render inside the read loop and confirm in the network panel that the response arrives in chunks. A single chunk named "the whole essay" means you are not streaming yet, even if the code mentions a stream. Fix that before you add a token counter or a fancy cursor. The feature is the early byte.
