# How the Streaming Chat API Provides Real-Time Responses with Context Retrieval in CodeWiki

> Discover how the streaming chat API delivers real-time AI responses by retrieving code context from your repository. Learn about its RAG pipeline using FAISS and Ollama.

- Repository: [Luong Quang Dung/codewiki](https://github.com/quangdungluong/codewiki)
- Tags: how-to-guide
- Published: 2026-02-16

---

**The streaming chat endpoint delivers AI-generated answers token-by-token while automatically retrieving relevant code context from your repository using a Retrieval-Augmented Generation (RAG) pipeline built with FAISS and Ollama embeddings.**

The `quangdungluong/codewiki` repository implements a sophisticated streaming architecture that combines real-time language model generation with intelligent context retrieval. This system enables developers to ask questions about specific codebases and receive immediate, contextually grounded responses without waiting for the entire answer to generate.

## Understanding the Streaming Chat Architecture

The core endpoint `POST /api/chat/stream` in [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py) orchestrates a four-stage pipeline that balances safety, retrieval accuracy, and real-time performance. Unlike standard request-response APIs, this implementation maintains an open HTTP connection using FastAPI's `StreamingResponse` with the `text/event-stream` MIME type, ensuring each token reaches the client immediately as Gemini produces it.

## The Four-Stage Pipeline for Real-Time Contextual Responses

### Stage 1: Request Validation and Token Safety

Before initiating retrieval or generation, the `stream_chat` function validates the incoming `ChatCompletionRequest` against strict safety criteria. The system verifies that the messages list is non-empty, ensures the last message originates from the user, and enforces an approximate 8,000-token limit on the final message using `utils/token_utils.count_tokens` from [`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py). This prevents context window overflow and ensures reliable streaming performance.

### Stage 2: Retrieval-Augmented Generation Setup

The RAG pipeline initializes by creating a `RAG` instance configured with the requested provider and model. The `prepare_retriever` method in [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py) constructs an Ollama-based embedder using the `nomic-embed-text` model to encode the user query. It then queries a pre-built FAISS index (managed by [`utils/localdb_manager.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/localdb_manager.py)) to retrieve the top 20 most relevant code documents. These documents are grouped by file path and concatenated into `context_text`, providing the language model with concrete code snippets to reference.

### Stage 3: Context-Aware Prompt Construction

The prompt assembly process in [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py) (lines 46-74) constructs a structured system prompt that includes repository metadata and usage guidelines. If conversation history exists, the `RAG.memory()` method injects previous exchanges. The retrieved context is wrapped between `<START_OF_CONTEXT>` and `<END_OF_CONTEXT>` tags, while the current user query is enclosed in `<query>...</query>` tags. This explicit delimitation ensures the Gemini model correctly distinguishes between historical conversation, retrieved code context, and the current question.

### Stage 4: Streaming Response Generation

The final stage invokes `genai.GenerativeModel.generate_content_async` with `stream=True`, creating an async generator that yields individual text chunks. The streaming loop (lines 85-92) iterates through these chunks, immediately yielding each fragment to FastAPI's `StreamingResponse`. The response headers explicitly disable buffering (`X-Accel-Buffering: no`), ensuring the client receives data instantly rather than waiting for the complete response. The front-end React application consumes this `text/event-stream`, appending each fragment to the chat interface to create a live typing effect.

## How Context Retrieval Works in the RAG Pipeline

The context retrieval mechanism operates through a sophisticated embedding and similarity search process. When `RAG.prepare_retriever` executes, it initializes the `nomic-embed-text` model via Ollama to create vector representations of the user query. The `LocalDBManager` in [`utils/localdb_manager.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/localdb_manager.py) maintains a FAISS index of the target repository's code documents, enabling millisecond-scale similarity searches.

The `FAISSRetriever` returns up to 20 relevant documents, which the system then groups by file path to eliminate redundancy and maintain code coherence. This grouped context is formatted with clear XML-like tags (`<START_OF_CONTEXT>`, `<END_OF_CONTEXT>`) to delineate retrieved information from conversation history and the current query. By injecting this concrete code context into the prompt before generation begins, the system ensures that responses reference actual repository implementations rather than hallucinating generic solutions.

## Client Integration Examples

### Using cURL for Real-Time Streaming

To test the streaming endpoint from the command line, use the `-N` flag to disable output buffering:

```bash
curl -N -X POST http://localhost:8001/api/chat/stream \
  -H "Content-Type: application/json" \
  -d '{
        "repo_url": "https://github.com/AsyncFuncAI/deepwiki-open",
        "messages": [
          {"role": "user", "content": "Explain how RAG works in this chat API"}
        ],
        "provider": "google",
        "model": "gemini-2.5-pro"
      }'

```

The `-N` flag ensures you see each token as it arrives, creating the real-time streaming effect.

### Python Synchronous Client with Requests

For Python applications, use the `requests` library with `stream=True` to consume the response iteratively:

```python
import requests

url = "http://localhost:8001/api/chat/stream"
payload = {
    "repo_url": "https://github.com/AsyncFuncAI/deepwiki-open",
    "messages": [{"role": "user", "content": "how backend process when user chat to repo, focus on how RAG works"}],
    "provider": "google",
    "model": "gemini-2.5-pro"
}

with requests.post(url, json=payload, stream=True) as r:
    for chunk in r.iter_content(decode_unicode=True):
        print(chunk, end='')

```

This pattern mirrors the test suite implementation and handles the `text/event-stream` response correctly.

### Async Python Client with HTTPX

For high-performance async applications, use `httpx` with async streaming:

```python
import httpx
import asyncio

async def stream_chat():
    async with httpx.AsyncClient() as client:
        resp = await client.post(
            "http://localhost:8001/api/chat/stream",
            json={
                "repo_url": "https://github.com/AsyncFuncAI/deepwiki-open",
                "messages": [{"role": "user", "content": "What files are most relevant to authentication?"}],
                "provider": "google",
                "model": "gemini-2.5-pro"
            },
            timeout=None,
        )
        async for chunk in resp.aiter_text():
            print(chunk, end='')

asyncio.run(stream_chat())

```

This async approach prevents blocking during I/O operations and is ideal for production applications handling multiple concurrent streams.

## Key Implementation Files

The streaming chat API with context retrieval is implemented across several specialized modules:

- **[`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py)** – Contains the main `stream_chat` endpoint that validates requests, orchestrates the RAG pipeline, constructs prompts, and manages the Gemini streaming response via `StreamingResponse`.

- **[`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py)** – Implements the `RAG` class with `prepare_retriever` and memory management methods. Handles Ollama embedding initialization and FAISS-based document retrieval.

- **[`api/models.py`](https://github.com/quangdungluong/codewiki/blob/main/api/models.py)** – Defines Pydantic schemas including `ChatCompletionRequest` for validating incoming chat requests with proper message structure and provider configuration.

- **[`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py)** – Provides `count_tokens` utility to enforce the 8,000-token safety limit on user messages before processing.

- **[`utils/localdb_manager.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/localdb_manager.py)** – Manages the FAISS vector index creation and persistence, enabling fast similarity searches across repository code documents.

These components work together to enable the real-time streaming experience while ensuring responses are grounded in actual repository context.

## Summary

- The streaming chat API delivers real-time responses through FastAPI's `StreamingResponse` with `text/event-stream`, pushing each token as Gemini generates it.
- Context retrieval uses a RAG pipeline with Ollama embeddings (`nomic-embed-text`) and FAISS similarity search to fetch up to 20 relevant code documents before generation begins.
- The system enforces safety through token counting ([`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py)) and request validation to prevent context window overflow.
- Retrieved context is formatted with XML-like tags (`<START_OF_CONTEXT>`, `<END_OF_CONTEXT>`) and combined with conversation history to create structured prompts for the LLM.
- Clients consume the stream using standard HTTP streaming techniques with disabled buffering to achieve the live typing effect.

## Frequently Asked Questions

### How does the streaming chat API handle large repository contexts without exceeding token limits?

The system implements a multi-layered approach to token management. First, [`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py) enforces an approximate 8,000-token limit on the user's final message through the `count_tokens` function. Second, the RAG pipeline in [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py) retrieves only the top 20 most relevant documents from the FAISS index, ensuring only pertinent context enters the prompt. Finally, the prompt construction logic in [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py) efficiently formats this context with minimal overhead using XML-like tags to delineate sections clearly.

### What embedding model does the context retrieval system use?

The RAG pipeline utilizes the `nomic-embed-text` model via Ollama for generating vector embeddings. This model is initialized in [`api/rag.py`](https://github.com/quangdungluong/codewiki/blob/main/api/rag.py) when the `RAG` class prepares its retriever component. The embeddings enable the FAISS-based similarity search to match user queries with relevant code snippets from the repository, providing the semantic foundation for the context retrieval mechanism.

### Can I use a different LLM provider with the streaming chat API?

Yes, the API supports configurable providers through the `ChatCompletionRequest` schema defined in [`api/models.py`](https://github.com/quangdungluong/codewiki/blob/main/api/models.py). While the implementation in [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py) specifically uses Google's Gemini via `genai.GenerativeModel.generate_content_async`, the request structure accepts a `provider` parameter (e.g., "google") and a `model` parameter (e.g., "gemini-2.5-pro"). This architecture allows for extension to support additional providers by implementing corresponding generation logic in the streaming handler.

### How does the client receive the stream without buffering delays?

The server explicitly disables buffering through HTTP response headers (`X-Accel-Buffering: no`) in the `StreamingResponse` configuration within [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py). Clients must also configure their HTTP libraries to disable buffering: cURL requires the `-N` flag, Python's `requests` library needs `stream=True`, and `httpx` requires async iteration over `aiter_text()`. This combination ensures each token generated by Gemini transmits immediately to the client, creating the real-time typing effect.