How the Streaming Chat API Provides Real-Time Responses with Context Retrieval in CodeWiki
The streaming chat endpoint delivers AI-generated answers token-by-token while automatically retrieving relevant code context from your repository using a Retrieval-Augmented Generation (RAG) pipeline built with FAISS and Ollama embeddings.
The quangdungluong/codewiki repository implements a sophisticated streaming architecture that combines real-time language model generation with intelligent context retrieval. This system enables developers to ask questions about specific codebases and receive immediate, contextually grounded responses without waiting for the entire answer to generate.
Understanding the Streaming Chat Architecture
The core endpoint POST /api/chat/stream in api/stream_chat.py orchestrates a four-stage pipeline that balances safety, retrieval accuracy, and real-time performance. Unlike standard request-response APIs, this implementation maintains an open HTTP connection using FastAPI's StreamingResponse with the text/event-stream MIME type, ensuring each token reaches the client immediately as Gemini produces it.
The Four-Stage Pipeline for Real-Time Contextual Responses
Stage 1: Request Validation and Token Safety
Before initiating retrieval or generation, the stream_chat function validates the incoming ChatCompletionRequest against strict safety criteria. The system verifies that the messages list is non-empty, ensures the last message originates from the user, and enforces an approximate 8,000-token limit on the final message using utils/token_utils.count_tokens from utils/token_utils.py. This prevents context window overflow and ensures reliable streaming performance.
Stage 2: Retrieval-Augmented Generation Setup
The RAG pipeline initializes by creating a RAG instance configured with the requested provider and model. The prepare_retriever method in api/rag.py constructs an Ollama-based embedder using the nomic-embed-text model to encode the user query. It then queries a pre-built FAISS index (managed by utils/localdb_manager.py) to retrieve the top 20 most relevant code documents. These documents are grouped by file path and concatenated into context_text, providing the language model with concrete code snippets to reference.
Stage 3: Context-Aware Prompt Construction
The prompt assembly process in api/stream_chat.py (lines 46-74) constructs a structured system prompt that includes repository metadata and usage guidelines. If conversation history exists, the RAG.memory() method injects previous exchanges. The retrieved context is wrapped between <START_OF_CONTEXT> and <END_OF_CONTEXT> tags, while the current user query is enclosed in <query>...</query> tags. This explicit delimitation ensures the Gemini model correctly distinguishes between historical conversation, retrieved code context, and the current question.
Stage 4: Streaming Response Generation
The final stage invokes genai.GenerativeModel.generate_content_async with stream=True, creating an async generator that yields individual text chunks. The streaming loop (lines 85-92) iterates through these chunks, immediately yielding each fragment to FastAPI's StreamingResponse. The response headers explicitly disable buffering (X-Accel-Buffering: no), ensuring the client receives data instantly rather than waiting for the complete response. The front-end React application consumes this text/event-stream, appending each fragment to the chat interface to create a live typing effect.
How Context Retrieval Works in the RAG Pipeline
The context retrieval mechanism operates through a sophisticated embedding and similarity search process. When RAG.prepare_retriever executes, it initializes the nomic-embed-text model via Ollama to create vector representations of the user query. The LocalDBManager in utils/localdb_manager.py maintains a FAISS index of the target repository's code documents, enabling millisecond-scale similarity searches.
The FAISSRetriever returns up to 20 relevant documents, which the system then groups by file path to eliminate redundancy and maintain code coherence. This grouped context is formatted with clear XML-like tags (<START_OF_CONTEXT>, <END_OF_CONTEXT>) to delineate retrieved information from conversation history and the current query. By injecting this concrete code context into the prompt before generation begins, the system ensures that responses reference actual repository implementations rather than hallucinating generic solutions.
Client Integration Examples
Using cURL for Real-Time Streaming
To test the streaming endpoint from the command line, use the -N flag to disable output buffering:
curl -N -X POST http://localhost:8001/api/chat/stream \
-H "Content-Type: application/json" \
-d '{
"repo_url": "https://github.com/AsyncFuncAI/deepwiki-open",
"messages": [
{"role": "user", "content": "Explain how RAG works in this chat API"}
],
"provider": "google",
"model": "gemini-2.5-pro"
}'
The -N flag ensures you see each token as it arrives, creating the real-time streaming effect.
Python Synchronous Client with Requests
For Python applications, use the requests library with stream=True to consume the response iteratively:
import requests
url = "http://localhost:8001/api/chat/stream"
payload = {
"repo_url": "https://github.com/AsyncFuncAI/deepwiki-open",
"messages": [{"role": "user", "content": "how backend process when user chat to repo, focus on how RAG works"}],
"provider": "google",
"model": "gemini-2.5-pro"
}
with requests.post(url, json=payload, stream=True) as r:
for chunk in r.iter_content(decode_unicode=True):
print(chunk, end='')
This pattern mirrors the test suite implementation and handles the text/event-stream response correctly.
Async Python Client with HTTPX
For high-performance async applications, use httpx with async streaming:
import httpx
import asyncio
async def stream_chat():
async with httpx.AsyncClient() as client:
resp = await client.post(
"http://localhost:8001/api/chat/stream",
json={
"repo_url": "https://github.com/AsyncFuncAI/deepwiki-open",
"messages": [{"role": "user", "content": "What files are most relevant to authentication?"}],
"provider": "google",
"model": "gemini-2.5-pro"
},
timeout=None,
)
async for chunk in resp.aiter_text():
print(chunk, end='')
asyncio.run(stream_chat())
This async approach prevents blocking during I/O operations and is ideal for production applications handling multiple concurrent streams.
Key Implementation Files
The streaming chat API with context retrieval is implemented across several specialized modules:
-
api/stream_chat.py– Contains the mainstream_chatendpoint that validates requests, orchestrates the RAG pipeline, constructs prompts, and manages the Gemini streaming response viaStreamingResponse. -
api/rag.py– Implements theRAGclass withprepare_retrieverand memory management methods. Handles Ollama embedding initialization and FAISS-based document retrieval. -
api/models.py– Defines Pydantic schemas includingChatCompletionRequestfor validating incoming chat requests with proper message structure and provider configuration. -
utils/token_utils.py– Providescount_tokensutility to enforce the 8,000-token safety limit on user messages before processing. -
utils/localdb_manager.py– Manages the FAISS vector index creation and persistence, enabling fast similarity searches across repository code documents.
These components work together to enable the real-time streaming experience while ensuring responses are grounded in actual repository context.
Summary
- The streaming chat API delivers real-time responses through FastAPI's
StreamingResponsewithtext/event-stream, pushing each token as Gemini generates it. - Context retrieval uses a RAG pipeline with Ollama embeddings (
nomic-embed-text) and FAISS similarity search to fetch up to 20 relevant code documents before generation begins. - The system enforces safety through token counting (
utils/token_utils.py) and request validation to prevent context window overflow. - Retrieved context is formatted with XML-like tags (
<START_OF_CONTEXT>,<END_OF_CONTEXT>) and combined with conversation history to create structured prompts for the LLM. - Clients consume the stream using standard HTTP streaming techniques with disabled buffering to achieve the live typing effect.
Frequently Asked Questions
How does the streaming chat API handle large repository contexts without exceeding token limits?
The system implements a multi-layered approach to token management. First, utils/token_utils.py enforces an approximate 8,000-token limit on the user's final message through the count_tokens function. Second, the RAG pipeline in api/rag.py retrieves only the top 20 most relevant documents from the FAISS index, ensuring only pertinent context enters the prompt. Finally, the prompt construction logic in api/stream_chat.py efficiently formats this context with minimal overhead using XML-like tags to delineate sections clearly.
What embedding model does the context retrieval system use?
The RAG pipeline utilizes the nomic-embed-text model via Ollama for generating vector embeddings. This model is initialized in api/rag.py when the RAG class prepares its retriever component. The embeddings enable the FAISS-based similarity search to match user queries with relevant code snippets from the repository, providing the semantic foundation for the context retrieval mechanism.
Can I use a different LLM provider with the streaming chat API?
Yes, the API supports configurable providers through the ChatCompletionRequest schema defined in api/models.py. While the implementation in api/stream_chat.py specifically uses Google's Gemini via genai.GenerativeModel.generate_content_async, the request structure accepts a provider parameter (e.g., "google") and a model parameter (e.g., "gemini-2.5-pro"). This architecture allows for extension to support additional providers by implementing corresponding generation logic in the streaming handler.
How does the client receive the stream without buffering delays?
The server explicitly disables buffering through HTTP response headers (X-Accel-Buffering: no) in the StreamingResponse configuration within api/stream_chat.py. Clients must also configure their HTTP libraries to disable buffering: cURL requires the -N flag, Python's requests library needs stream=True, and httpx requires async iteration over aiter_text(). This combination ensures each token generated by Gemini transmits immediately to the client, creating the real-time typing effect.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →