# How to Enable and Implement Streaming Responses for Chat and Completion Endpoints in PrivateGPT

> Enable streaming responses in PrivateGPT for chat and completion endpoints. Learn how to use SSE for incremental token delivery instead of full JSON payloads. Optimize your applications today.

- Repository: [Zylon/private-gpt](https://github.com/zylon-ai/private-gpt)
- Tags: how-to-guide
- Published: 2026-03-06

---

**To enable streaming responses in PrivateGPT, set `"stream": true` in your request body to trigger the SSE (Server-Sent Events) code path in the FastAPI routers, which returns tokens incrementally instead of a full JSON payload.**

PrivateGPT (zylon-ai/private-gpt) provides OpenAI-compatible chat and completion endpoints that support real-time token streaming via Server-Sent Events (SSE). This implementation allows clients to receive LLM output incrementally as it is generated, significantly reducing latency perception in interactive applications.

## How Streaming Is Triggered

Streaming mode is activated through a conditional code path in both the chat and completion FastAPI routers. The client must explicitly request streaming by including the **boolean field `stream`** in the request body.

In [`private_gpt/server/chat/chat_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_router.py), the `stream` field is defined in `ChatBody` (lines 19-25). Similarly, the completion router defines it in `CompletionsBody.stream` within [`private_gpt/server/completions/completions_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/completions/completions_router.py) (lines 16-23).

When `body.stream` evaluates to `true`, the router bypasses the standard JSON response and instead returns a `StreamingResponse` with `media_type="text/event-stream"`. This occurs in the chat router at lines 94-101 and in the completion router at lines 72-80.

## The Service Layer Implementation

The streaming logic delegates to **`ChatService.stream_chat`** defined in [`private_gpt/server/chat/chat_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chat/chat_service.py) (lines 49-84). This method constructs a **`CompletionGen`** object containing a generator expression (`response_gen`) that yields individual tokens from the underlying LLM engine as they become available.

The `stream_chat` method handles both the LLM interaction and optional source document retrieval. When `include_sources` is enabled alongside streaming, the service attaches relevant `Chunk` objects (defined in [`private_gpt/server/chunks/chunks_service.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/chunks/chunks_service.py)) to the stream.

## SSE Conversion and OpenAI Compatibility

Raw token generators are converted to OpenAI-compatible Server-Sent Events through the **`to_openai_sse_stream`** utility function located in [`private_gpt/open_ai/openai_models.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/open_ai/openai_models.py) (lines 12-22).

This function iterates over the generator and formats each chunk according to the OpenAI streaming specification:

- Each token is wrapped in a JSON payload containing `id`, `object`, and `choices` fields
- The payload is prefixed with `data: ` to conform to SSE standards
- A final `data: [DONE]` marker signals stream completion

The resulting iterator is wrapped in FastAPI's `StreamingResponse`, ensuring proper HTTP headers for event streaming.

## Practical Implementation Examples

### Streaming Chat Completions with cURL

Use the `-N` flag to disable buffering and see tokens arrive in real-time:

```bash
curl -N -X POST "http://localhost:8000/v1/chat/completions" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "messages": [
          {"role": "system", "content": "You are a helpful assistant."},
          {"role": "user",   "content": "Explain quantum computing in one sentence."}
        ],
        "stream": true,
        "include_sources": false
      }'

```

Each line arrives as `data: {"id":"...","object":"chat.completion.chunk","choices":[{"delta":{"content":"token"}}]}` followed by `data: [DONE]`.

### Streaming Classic Completions

The completion endpoint follows identical patterns in [`private_gpt/server/completions/completions_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/server/completions/completions_router.py):

```bash
curl -N -X POST "http://localhost:8000/v1/completions" \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "prompt": "Write a haiku about sunrise.",
        "stream": true,
        "include_sources": true
      }'

```

When `include_sources` is `true`, the `sources` field appears in each chunk's `choices[0]` object, allowing you to stream both generated text and retrieved document chunks simultaneously.

### Consuming Streams in Python

For programmatic access, use `requests` with `stream=True`:

```python
import requests

url = "http://localhost:8000/v1/chat/completions"
headers = {"Authorization": "Bearer YOUR_TOKEN"}
payload = {
    "messages": [
        {"role": "system", "content": "You are a poet."},
        {"role": "user",   "content": "Compose a limerick about cats."}
    ],
    "stream": True,
}

with requests.post(url, json=payload, headers=headers, stream=True) as resp:
    for line in resp.iter_lines():
        if line:
            print(line.decode())

```

This prints each SSE line immediately as the LLM generates tokens, enabling real-time display in user interfaces.

## Summary

- **Enable streaming** by setting `"stream": true` in request bodies for both chat and completion endpoints
- **Architecture**: FastAPI routers check the `stream` boolean and delegate to `ChatService.stream_chat`, which returns a `CompletionGen` generator
- **SSE formatting**: The `to_openai_sse_stream` function in [`openai_models.py`](https://github.com/zylon-ai/private-gpt/blob/main/openai_models.py) converts tokens to OpenAI-compatible Server-Sent Events with proper `data:` prefixes and `[DONE]` termination
- **Source streaming**: Include `"include_sources": true` to stream retrieved document chunks alongside generated text
- **No configuration required**: The streaming infrastructure is fully implemented in the codebase; only the client request parameter needs adjustment

## Frequently Asked Questions

### How do I enable streaming responses in PrivateGPT?

Set the `stream` parameter to `true` in your JSON request body when calling either `/v1/chat/completions` or `/v1/completions`. This triggers the conditional logic in the FastAPI routers (lines 94-101 in [`chat_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/chat_router.py) and lines 72-80 in [`completions_router.py`](https://github.com/zylon-ai/private-gpt/blob/main/completions_router.py)) to return a `StreamingResponse` instead of a complete JSON object.

### What format do streaming responses use?

PrivateGPT implements OpenAI-compatible Server-Sent Events (SSE). Each token is delivered as a line prefixed with `data: ` containing a JSON chunk with `id`, `object`, and `choices` fields. The stream terminates with `data: [DONE]`. This format is generated by the `to_openai_sse_stream` function in [`private_gpt/open_ai/openai_models.py`](https://github.com/zylon-ai/private-gpt/blob/main/private_gpt/open_ai/openai_models.py).

### Can I stream source documents alongside the generated text?

Yes. Set both `"stream": true` and `"include_sources": true` in your request. The service layer attaches `Chunk` objects to each streamed segment, allowing you to receive retrieved context chunks interleaved with or alongside the LLM's token generation.

### Is there a performance benefit to using streaming?

Streaming reduces time-to-first-byte significantly, as the client receives the first token immediately rather than waiting for the complete response. However, total generation time remains identical to non-streaming mode. The primary benefit is improved perceived latency and responsiveness in interactive applications.