How to Enable and Implement Streaming Responses for Chat and Completion Endpoints in PrivateGPT
To enable streaming responses in PrivateGPT, set "stream": true in your request body to trigger the SSE (Server-Sent Events) code path in the FastAPI routers, which returns tokens incrementally instead of a full JSON payload.
PrivateGPT (zylon-ai/private-gpt) provides OpenAI-compatible chat and completion endpoints that support real-time token streaming via Server-Sent Events (SSE). This implementation allows clients to receive LLM output incrementally as it is generated, significantly reducing latency perception in interactive applications.
How Streaming Is Triggered
Streaming mode is activated through a conditional code path in both the chat and completion FastAPI routers. The client must explicitly request streaming by including the boolean field stream in the request body.
In private_gpt/server/chat/chat_router.py, the stream field is defined in ChatBody (lines 19-25). Similarly, the completion router defines it in CompletionsBody.stream within private_gpt/server/completions/completions_router.py (lines 16-23).
When body.stream evaluates to true, the router bypasses the standard JSON response and instead returns a StreamingResponse with media_type="text/event-stream". This occurs in the chat router at lines 94-101 and in the completion router at lines 72-80.
The Service Layer Implementation
The streaming logic delegates to ChatService.stream_chat defined in private_gpt/server/chat/chat_service.py (lines 49-84). This method constructs a CompletionGen object containing a generator expression (response_gen) that yields individual tokens from the underlying LLM engine as they become available.
The stream_chat method handles both the LLM interaction and optional source document retrieval. When include_sources is enabled alongside streaming, the service attaches relevant Chunk objects (defined in private_gpt/server/chunks/chunks_service.py) to the stream.
SSE Conversion and OpenAI Compatibility
Raw token generators are converted to OpenAI-compatible Server-Sent Events through the to_openai_sse_stream utility function located in private_gpt/open_ai/openai_models.py (lines 12-22).
This function iterates over the generator and formats each chunk according to the OpenAI streaming specification:
- Each token is wrapped in a JSON payload containing
id,object, andchoicesfields - The payload is prefixed with
data:to conform to SSE standards - A final
data: [DONE]marker signals stream completion
The resulting iterator is wrapped in FastAPI's StreamingResponse, ensuring proper HTTP headers for event streaming.
Practical Implementation Examples
Streaming Chat Completions with cURL
Use the -N flag to disable buffering and see tokens arrive in real-time:
curl -N -X POST "http://localhost:8000/v1/chat/completions" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in one sentence."}
],
"stream": true,
"include_sources": false
}'
Each line arrives as data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"token"}}]} followed by data: [DONE].
Streaming Classic Completions
The completion endpoint follows identical patterns in private_gpt/server/completions/completions_router.py:
curl -N -X POST "http://localhost:8000/v1/completions" \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"prompt": "Write a haiku about sunrise.",
"stream": true,
"include_sources": true
}'
When include_sources is true, the sources field appears in each chunk's choices[0] object, allowing you to stream both generated text and retrieved document chunks simultaneously.
Consuming Streams in Python
For programmatic access, use requests with stream=True:
import requests
url = "http://localhost:8000/v1/chat/completions"
headers = {"Authorization": "Bearer YOUR_TOKEN"}
payload = {
"messages": [
{"role": "system", "content": "You are a poet."},
{"role": "user", "content": "Compose a limerick about cats."}
],
"stream": True,
}
with requests.post(url, json=payload, headers=headers, stream=True) as resp:
for line in resp.iter_lines():
if line:
print(line.decode())
This prints each SSE line immediately as the LLM generates tokens, enabling real-time display in user interfaces.
Summary
- Enable streaming by setting
"stream": truein request bodies for both chat and completion endpoints - Architecture: FastAPI routers check the
streamboolean and delegate toChatService.stream_chat, which returns aCompletionGengenerator - SSE formatting: The
to_openai_sse_streamfunction inopenai_models.pyconverts tokens to OpenAI-compatible Server-Sent Events with properdata:prefixes and[DONE]termination - Source streaming: Include
"include_sources": trueto stream retrieved document chunks alongside generated text - No configuration required: The streaming infrastructure is fully implemented in the codebase; only the client request parameter needs adjustment
Frequently Asked Questions
How do I enable streaming responses in PrivateGPT?
Set the stream parameter to true in your JSON request body when calling either /v1/chat/completions or /v1/completions. This triggers the conditional logic in the FastAPI routers (lines 94-101 in chat_router.py and lines 72-80 in completions_router.py) to return a StreamingResponse instead of a complete JSON object.
What format do streaming responses use?
PrivateGPT implements OpenAI-compatible Server-Sent Events (SSE). Each token is delivered as a line prefixed with data: containing a JSON chunk with id, object, and choices fields. The stream terminates with data: [DONE]. This format is generated by the to_openai_sse_stream function in private_gpt/open_ai/openai_models.py.
Can I stream source documents alongside the generated text?
Yes. Set both "stream": true and "include_sources": true in your request. The service layer attaches Chunk objects to each streamed segment, allowing you to receive retrieved context chunks interleaved with or alongside the LLM's token generation.
Is there a performance benefit to using streaming?
Streaming reduces time-to-first-byte significantly, as the client receives the first token immediately rather than waiting for the complete response. However, total generation time remains identical to non-streaming mode. The primary benefit is improved perceived latency and responsiveness in interactive applications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →