# How CodeWiki Manages Token Counting and Context Window in Chat Conversations

> Discover how CodeWiki manages token counting and context window with a three-layer defense: tiktoken, message limits, and document thresholds to optimize chat performance.

- Repository: [Luong Quang Dung/codewiki](https://github.com/quangdungluong/codewiki)
- Tags: internals
- Published: 2026-02-16

---

**CodeWiki uses a three-layer defense strategy that combines tiktoken-based counting, an 8,000-token limit on user messages, and document-specific thresholds to keep chat interactions within the model's context window.**

CodeWiki, an open-source retrieval-augmented generation (RAG) chatbot built by `quangdungluong/codewiki`, must carefully manage token counting and context window constraints to prevent overwhelming the underlying language model. The system implements coordinated checks across the request pipeline to ensure that both user prompts and retrieved documents fit within the available token budget.

## Token Counting Implementation in CodeWiki

### The count_tokens Utility Function

At the foundation of CodeWiki's token management lies a lightweight wrapper around OpenAI's `tiktoken` library. Located in [`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py), the `count_tokens` function initializes the `cl100k_base` encoding and returns the exact token count for any input string.

```python

# utils/token_utils.py

import tiktoken

def count_tokens(text: str) -> int:
    encoding = tiktoken.get_encoding("cl100k_base")
    return len(encoding.encode(text))

```

This utility is imported throughout the codebase wherever token-budget decisions are required, ensuring consistent counting methodology across the chat endpoint and document processing pipeline.

## Guarding the User Prompt Against Context Window Limits

### The 8,000 Token Threshold

When a chat request arrives at the `/stream` endpoint in [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py), the system immediately inspects the last user message to protect the model's input context window. If the message exceeds **8,000 tokens**, CodeWiki flags the request as `input_too_large` and aborts the context-building phase.

```python

# api/stream_chat.py (lines 30-38)

tokens = count_tokens(last_message.content)
if tokens > 8000:
    input_too_large = True   # aborts context-building

```

This check prevents the model from receiving oversized prompts that would exceed the roughly 8K-token limit allocated for user-provided content. When triggered, the system answers the query without retrieval augmentation, ensuring the conversation remains within the model's operational constraints.

## Filtering Retrieved Documents by Token Size

### Code Files vs. Documentation Limits

Before retrieved documents become part of the prompt, CodeWiki applies strict size filters in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py). The system distinguishes between code files and plain-text documentation, applying different multipliers of the `MAX_EMBEDDING_TOKEN` constant (defined as **8,192** in [`utils/constants.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/constants.py)).

For **code files**, the threshold is **10 × MAX_EMBEDDING_TOKEN** (approximately 81,920 tokens). Files exceeding this limit are dropped from the context. For **documentation**, the stricter **MAX_EMBEDDING_TOKEN** limit (8,192 tokens) applies.

```python

# utils/document_pipeline.py (lines 27-34 & 65-71)

token_count = count_tokens(content)
if token_count > MAX_EMBEDDING_TOKEN * 10:   # code files

    continue
...
if token_count > MAX_EMBEDDING_TOKEN:        # doc files

    continue

```

After filtering, remaining documents are grouped by file path and concatenated into a single context block wrapped in `<START_OF_CONTEXT>` and `<END_OF_CONTEXT>` markers. If the request was previously flagged as `input_too_large`, this entire context block is omitted.

## Summary

- **Centralized token counting**: CodeWiki uses a single `count_tokens` function in [`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py) wrapping `tiktoken` with the `cl100k_base` encoding to ensure consistent measurement across the application.
- **User message protection**: The chat endpoint in [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py) rejects or bypasses retrieval when user messages exceed 8,000 tokens, safeguarding the model's input context window.
- **Document size filtering**: The pipeline in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) applies tiered limits—81,920 tokens for code files and 8,192 tokens for documentation—based on the `MAX_EMBEDDING_TOKEN` constant from [`utils/constants.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/constants.py).
- **Graceful degradation**: When token limits are breached, CodeWiki falls back to non-retrieval responses rather than failing, ensuring continuous service availability.

## Frequently Asked Questions

### How does CodeWiki count tokens for chat messages?

CodeWiki counts tokens using a dedicated utility function `count_tokens` located in [`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py). This function initializes the `cl100k_base` encoding from OpenAI's `tiktoken` library and returns the length of the encoded token list. All components of the chat pipeline import this function to ensure consistent token measurement.

### What happens when a user message exceeds the token limit?

When a user message exceeds 8,000 tokens, the chat endpoint in [`api/stream_chat.py`](https://github.com/quangdungluong/codewiki/blob/main/api/stream_chat.py) sets an `input_too_large` flag and skips the document retrieval phase. The model then generates a response based solely on the user's query without augmentation from the codebase context. This prevents the model from receiving an oversized prompt that would exceed its context window capacity.

### Why does CodeWiki apply different token limits to code files versus documentation?

CodeWiki applies a higher token limit to code files (81,920 tokens, or 10× `MAX_EMBEDDING_TOKEN`) compared to documentation (8,192 tokens) because source code files often contain substantial implementation details necessary for accurate responses. The `MAX_EMBEDDING_TOKEN` constant defined in [`utils/constants.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/constants.py) serves as the baseline for these calculations, allowing the system to retain large codebases in context while keeping plain-text documentation concise.