# Document Processing Pipeline for Embedding Code in Codewiki: A 5-Stage Technical Breakdown

> Explore quangdungluongcodewiki's 5-stage document processing pipeline for embedding code. Learn how source code is cloned filtered chunked embedded and cached for RAG.

- Repository: [Luong Quang Dung/codewiki](https://github.com/quangdungluong/codewiki)
- Tags: internals
- Published: 2026-02-16

---

**The codewiki repository implements a five-stage document processing pipeline that clones source code repositories, filters and chunks files using tiktoken limits, generates embeddings via Ollama's nomic-embed-text model, and persists vectors to a local cache for retrieval-augmented generation.**

The **document processing pipeline for embedding code** is the core infrastructure that powers codewiki's ability to transform raw repositories into searchable vector indexes. This pipeline orchestrates repository acquisition, intelligent document parsing, token-aware chunking, and local embedding generation using the AdalFlow framework. Understanding each stage is essential for developers looking to customize the ingestion flow or debug embedding quality issues.

## Stage 1: Repository Acquisition

The pipeline begins by fetching the target repository into a local workspace. The `RepoDownloader` class in [`utils/localdb_manager.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/localdb_manager.py) handles the cloning or fetching process, ensuring the codebase is available for local analysis. This component is invoked within the `_create_repo` method of the `LocalDBManager`, which coordinates the download before passing control to the document reader.

## Stage 2: Recursive Document Reading

Once the repository is local, the `RecursiveDocumentReader` class in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) traverses the directory tree via the `read_documents()` method. This stage applies several critical filters to ensure only processable content enters the embedding queue.

### File Filtering and Token Limits

The reader enforces strict token budgets using `tiktoken` (implemented in [`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py)). As defined in [`utils/constants.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/constants.py), the `MAX_EMBEDDING_TOKEN` is set to **8192**. Code files are permitted up to `MAX_EMBEDDING_TOKEN * 10` tokens before rejection, while documentation files must fit within the base limit. Files are also filtered by extension and exclusion lists (e.g., `.git`, `tests` directories).

### Metadata Enrichment

Each qualifying file is wrapped as an `adalflow.core.types.Document` object enriched with metadata including `file_path`, `is_implementation` (determined by path heuristics), `type`, `is_code`, `title`, and `token_count`. This metadata enables downstream retrievers to bias results toward implementation files versus test files.

## Stage 3: Chunking and Embedding Preparation

The `DocumentTransformer` class in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) orchestrates the transformation logic through its `_prepare_data_pipeline()` method. This stage utilizes **AdalFlow's** `Sequential` component to chain operations:

1. **Splitting**: Documents are divided into overlapping chunks of **500 tokens** with a **100-token overlap** to preserve context across boundaries.
2. **Embedder Initialization**: An Ollama embedder is instantiated using the `nomic-embed-text` model, wrapped within the AdalFlow `Embedder` interface.

This configuration ensures that large source files are broken into semantically coherent segments suitable for vector search while maintaining the locality of reference necessary for code understanding.

## Stage 4: Per-Document Embedding

Due to limitations in the Ollama client (which does not support batch embedding calls in this implementation), each chunk is processed individually. The `OllamaDocumentProcessor` class in [`utils/ollama_embedder.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/ollama_embedder.py) handles this via its `__call__()` method, which sends each text chunk to the local Ollama instance and attaches the returned vector to the document's `vector` attribute.

This stage is the computational bottleneck of the pipeline, as sequential API calls to the embedding model are slower than batch operations. However, it ensures compatibility with local Ollama deployments without requiring GPU batching infrastructure.

## Stage 5: Persistence to LocalDB

The final stage persists the transformed documents to disk to avoid re-embedding on subsequent runs. The `DocumentTransformer.transform_and_save()` method (coordinated through `LocalDBManager`) serializes the list of documents—complete with chunks and vectors—to a pickle file within the local `LocalDB` cache directory defined in [`utils/constants.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/constants.py).

When the system initializes, it checks this cache first. If a valid index exists for the requested repository, the pipeline skips directly to retrieval, significantly reducing startup latency for frequently accessed codebases.

## Complete Pipeline Implementation Example

The following example demonstrates how to invoke the entire document processing pipeline using the `LocalDBManager` orchestrator:

```python
from utils.localdb_manager import LocalDBManager

# Initialise manager

db_manager = LocalDBManager()

# Prepare the index (downloads repo, reads files, chunks, embeds, persists)

documents = db_manager.prepare_retriever(
    repo_url="https://github.com/quangdungluong/codewiki.git",
    access_token=None,               # optional GitHub token

    excluded_dirs=[".git", "tests"], # example exclusions

)

print(f"Embedded {len(documents)} chunks ready for retrieval")

# Each `documents[i]` now has `vector` (the embedding) and metadata

```

This high-level interface abstracts the five stages, providing a single entry point for repository ingestion while exposing configuration options for authentication and directory exclusion.

## Key Configuration Constants

The pipeline behavior is governed by constants defined in [`utils/constants.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/constants.py):

- **`MAX_EMBEDDING_TOKEN = 8192`**: The base token limit for embedding context windows.
- **Code file limit**: `MAX_EMBEDDING_TOKEN * 10` (81,920 tokens), allowing large implementation files to be processed.
- **Documentation limit**: Strict `MAX_EMBEDDING_TOKEN` (8,192 tokens) for non-code files.
- **Chunking parameters**: 500-token chunks with 100-token overlap, balancing granularity and context preservation.

Token counting is performed using `tiktoken` via [`utils/token_utils.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/token_utils.py), ensuring accurate OpenAI-compatible token estimates regardless of file encoding.

## Summary

- The **document processing pipeline for embedding code** in codewiki consists of five sequential stages: repository acquisition, recursive document reading with token filtering, chunking and embedding preparation, per-document vector generation via Ollama, and persistence to a local cache.
- Core implementation files include [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) for reading and chunking, [`utils/ollama_embedder.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/ollama_embedder.py) for embedding generation, and [`utils/localdb_manager.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/localdb_manager.py) for orchestration and caching.
- The pipeline enforces strict token limits (8,192 base limit, 81,920 for code files) using `tiktoken` and creates overlapping 500-token chunks to maintain semantic context.
- Due to Ollama client constraints, embeddings are generated sequentially rather than in batches, with results stored as pickle files for reuse via `LocalDBManager`.

## Frequently Asked Questions

### How does codewiki handle large code files that exceed standard token limits?

The pipeline applies differentiated token limits based on file type. According to [`utils/constants.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/constants.py), code files are permitted up to **81,920 tokens** (`MAX_EMBEDDING_TOKEN * 10`), while documentation files are strictly capped at **8,192 tokens**. Files exceeding these thresholds are filtered out during the `RecursiveDocumentReader.read_documents()` stage in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) to prevent context window overflow during embedding.

### What chunking strategy does the document processing pipeline use for embedding code?

The `DocumentTransformer` class in [`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py) splits documents into **500-token chunks with a 100-token overlap**. This overlapping window strategy ensures that code context—such as function definitions spanning multiple blocks—is preserved across chunk boundaries. The chunking is implemented within the `_prepare_data_pipeline()` method using AdalFlow's sequential transformation components.

### Why does codewiki embed documents individually rather than in batches?

The current implementation uses the `OllamaDocumentProcessor` class in [`utils/ollama_embedder.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/ollama_embedder.py), which processes each chunk individually via its `__call__()` method. This sequential approach is necessitated by the Ollama client's lack of batch embedding support in this specific integration. While this creates a processing bottleneck compared to batch APIs, it ensures compatibility with local Ollama deployments without requiring additional GPU batching infrastructure.

### How does the pipeline persist embeddings to avoid reprocessing repositories?

After embedding generation, the `DocumentTransformer.transform_and_save()` method serializes the complete list of documents—including their chunks and vector embeddings—to a pickle file managed by `LocalDBManager` in [`utils/localdb_manager.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/localdb_manager.py). When `prepare_retriever()` is called subsequently, the system checks for existing cached indices and loads them directly, skipping the download, chunking, and embedding stages entirely.