How CodeWiki Manages Token Counting and Context Window in Chat Conversations

CodeWiki uses a three-layer defense strategy that combines tiktoken-based counting, an 8,000-token limit on user messages, and document-specific thresholds to keep chat interactions within the model's context window.

CodeWiki, an open-source retrieval-augmented generation (RAG) chatbot built by quangdungluong/codewiki, must carefully manage token counting and context window constraints to prevent overwhelming the underlying language model. The system implements coordinated checks across the request pipeline to ensure that both user prompts and retrieved documents fit within the available token budget.

Token Counting Implementation in CodeWiki

The count_tokens Utility Function

At the foundation of CodeWiki's token management lies a lightweight wrapper around OpenAI's tiktoken library. Located in utils/token_utils.py, the count_tokens function initializes the cl100k_base encoding and returns the exact token count for any input string.


# utils/token_utils.py

import tiktoken

def count_tokens(text: str) -> int:
    encoding = tiktoken.get_encoding("cl100k_base")
    return len(encoding.encode(text))

This utility is imported throughout the codebase wherever token-budget decisions are required, ensuring consistent counting methodology across the chat endpoint and document processing pipeline.

Guarding the User Prompt Against Context Window Limits

The 8,000 Token Threshold

When a chat request arrives at the /stream endpoint in api/stream_chat.py, the system immediately inspects the last user message to protect the model's input context window. If the message exceeds 8,000 tokens, CodeWiki flags the request as input_too_large and aborts the context-building phase.


# api/stream_chat.py (lines 30-38)

tokens = count_tokens(last_message.content)
if tokens > 8000:
    input_too_large = True   # aborts context-building

This check prevents the model from receiving oversized prompts that would exceed the roughly 8K-token limit allocated for user-provided content. When triggered, the system answers the query without retrieval augmentation, ensuring the conversation remains within the model's operational constraints.

Filtering Retrieved Documents by Token Size

Code Files vs. Documentation Limits

Before retrieved documents become part of the prompt, CodeWiki applies strict size filters in utils/document_pipeline.py. The system distinguishes between code files and plain-text documentation, applying different multipliers of the MAX_EMBEDDING_TOKEN constant (defined as 8,192 in utils/constants.py).

For code files, the threshold is 10 × MAX_EMBEDDING_TOKEN (approximately 81,920 tokens). Files exceeding this limit are dropped from the context. For documentation, the stricter MAX_EMBEDDING_TOKEN limit (8,192 tokens) applies.


# utils/document_pipeline.py (lines 27-34 & 65-71)

token_count = count_tokens(content)
if token_count > MAX_EMBEDDING_TOKEN * 10:   # code files

    continue
...
if token_count > MAX_EMBEDDING_TOKEN:        # doc files

    continue

After filtering, remaining documents are grouped by file path and concatenated into a single context block wrapped in <START_OF_CONTEXT> and <END_OF_CONTEXT> markers. If the request was previously flagged as input_too_large, this entire context block is omitted.

Summary

  • Centralized token counting: CodeWiki uses a single count_tokens function in utils/token_utils.py wrapping tiktoken with the cl100k_base encoding to ensure consistent measurement across the application.
  • User message protection: The chat endpoint in api/stream_chat.py rejects or bypasses retrieval when user messages exceed 8,000 tokens, safeguarding the model's input context window.
  • Document size filtering: The pipeline in utils/document_pipeline.py applies tiered limits—81,920 tokens for code files and 8,192 tokens for documentation—based on the MAX_EMBEDDING_TOKEN constant from utils/constants.py.
  • Graceful degradation: When token limits are breached, CodeWiki falls back to non-retrieval responses rather than failing, ensuring continuous service availability.

Frequently Asked Questions

How does CodeWiki count tokens for chat messages?

CodeWiki counts tokens using a dedicated utility function count_tokens located in utils/token_utils.py. This function initializes the cl100k_base encoding from OpenAI's tiktoken library and returns the length of the encoded token list. All components of the chat pipeline import this function to ensure consistent token measurement.

What happens when a user message exceeds the token limit?

When a user message exceeds 8,000 tokens, the chat endpoint in api/stream_chat.py sets an input_too_large flag and skips the document retrieval phase. The model then generates a response based solely on the user's query without augmentation from the codebase context. This prevents the model from receiving an oversized prompt that would exceed its context window capacity.

Why does CodeWiki apply different token limits to code files versus documentation?

CodeWiki applies a higher token limit to code files (81,920 tokens, or 10× MAX_EMBEDDING_TOKEN) compared to documentation (8,192 tokens) because source code files often contain substantial implementation details necessary for accurate responses. The MAX_EMBEDDING_TOKEN constant defined in utils/constants.py serves as the baseline for these calculations, allowing the system to retain large codebases in context while keeping plain-text documentation concise.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →