# Wiki Generation Pipeline and Redis Task Queuing in CodeWiki: A Technical Deep Dive

> Explore the CodeWiki generation pipeline a multi-stage workflow that converts code to wikis. Learn how it uses Redis task queuing for efficient asynchronous processing and worker coordination.

- Repository: [Luong Quang Dung/codewiki](https://github.com/quangdungluong/codewiki)
- Tags: deep-dive
- Published: 2026-02-16

---

**The wiki generation pipeline is an asynchronous, multi-stage workflow that transforms code repositories into structured technical wikis, using Redis as a lightweight state store to track progress and coordinate concurrent workers without requiring a dedicated message broker.**

The `quangdungluong/codewiki` project implements a documentation engine that converts source code into AI-generated wikis through a sophisticated **wiki generation pipeline**. This system leverages **FastAPI** endpoints and **Redis task queuing** to orchestrate repository analysis, LLM-powered content creation, and real-time status delivery.

## Seven Stages of the Wiki Generation Pipeline

### Stage 1: Task Initialization

When a client POSTs a `WikiTaskRequest` to **`/api/wiki/generate`**, the system generates a UUID and writes an initial task entry to Redis with status `started`. FastAPI’s `BackgroundTasks` then hands the request to the processing engine. This logic resides in [[`api/wiki.py`](https://github.com/quangdungluong/codewiki/blob/main/api/wiki.py)](https://github.com/quangdungluong/codewiki/blob/master/api/wiki.py) lines 33-48.

### Stage 2: Repository Structure Fetching

The `RepositoryStructureFetcher` contacts the appropriate source—whether a local API or the GitHub API—to build a complete file tree and README snapshot. This phase executes within [[`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py)](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 68-115.

### Stage 3: Wiki Structure Determination

The fetched repository data is fed to the Gemini model (`gemini-2.5-pro`), which returns an **XML** description of pages, sections, and their hierarchical relationships. The system parses this XML into internal Pydantic models: `WikiStructure`, `WikiPage`, and `WikiSection`. See [[`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py)](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 180-388.

### Stage 4: Concurrent Page Content Generation

For every page defined in the structure, background workers invoke the Gemini model with detailed prompts enforcing specific Markdown formats like `<details>` blocks, Mermaid diagrams, and tables. By default, **three workers** run concurrently to generate content in parallel. This concurrency logic appears in [[`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py)](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 500-618.

### Stage 5: Real-Time Status Persistence

After each significant step—fetching, structure building, or page completion—the pipeline calls `RedisTasks.update_task` to mutate the JSON stored under the task UUID. This JSON contains `status`, `message`, optional `error` details, partial `WikiCacheData`, and a `progress` list tracking currently processing page IDs. The update methods are located in [[`api/wiki.py`](https://github.com/quangdungluong/codewiki/blob/main/api/wiki.py)](https://github.com/quangdungluong/codewiki/blob/master/api/wiki.py) lines 15-31, 76-84, and 122-130.

### Stage 6: Client Polling Interface

Clients monitor progress by querying **`/api/wiki/status/{task_id}`**, which reads the latest task entry from Redis via `RedisTasks.get_task` and returns the current JSON payload. This endpoint is implemented in [[`api/wiki.py`](https://github.com/quangdungluong/codewiki/blob/main/api/wiki.py)](https://github.com/quangdungluong/codewiki/blob/master/api/wiki.py) lines 65-71.

### Stage 7: Cache Persistence

Upon successful completion, the assembled `WikiStructure` and generated page contents are POSTed to the cache service at `/api/wiki_cache`. This final persistence step is handled in [[`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py)](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 438-447.

## Redis Task State Management

The system uses **Redis** not as a traditional message queue, but as an atomic state store. The `RedisTasks` class in [[`utils/redis_tasks.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/redis_tasks.py)](https://github.com/quangdungluong/codewiki/blob/master/utils/redis_tasks.py) provides a thin wrapper around `redis-py`, implementing `add_task`, `get_task`, `update_task`, and `delete_task` operations.

Each task is stored as a JSON string keyed by its UUID. Unlike systems using separate queue data structures, the task record itself acts as the **single source of truth** for status and progress. Because Redis supports atomic `SET` and `GET` operations, multiple concurrent workers can safely update task state without race conditions. The background workers write updates (e.g., `status="processing"`, `progress=[...]`) while the API endpoints read the current values.

## Core Implementation Files

| File | Role |
|------|------|
| **[`api/wiki.py`](https://github.com/quangdungluong/codewiki/blob/main/api/wiki.py)** | FastAPI router handling task creation at `/api/wiki/generate`, status polling at `/api/wiki/status/{task_id}`, and background task orchestration. |
| **[`utils/redis_tasks.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/redis_tasks.py)** | Redis wrapper managing atomic task state operations (`add_task`, `get_task`, `update_task`, `delete_task`). |
| **[`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py)** | Core engine fetching repository data, invoking the LLM for structure generation, managing concurrent page workers, and updating Redis state throughout execution. |
| **[`utils/document_pipeline.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/document_pipeline.py)** | Provides `RecursiveDocumentReader` and `DocumentTransformer` for source file processing used during content generation. |
| **[`utils/models.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/models.py)** | Pydantic models (`WikiPage`, `WikiSection`, `WikiStructure`, `WikiCacheData`) for data validation and serialization across the pipeline. |

## Code Examples: Interacting with the Pipeline

### Triggering a New Wiki Generation

Use `httpx` to submit a repository for processing:

```python
import httpx
import uuid

payload = {
    "repo_info": {"type": "web"},
    "repo_url": "https://github.com/quangdungluong/codewiki",
    "owner": "quangdungluong",
    "repo": "codewiki",
    "token": None,
}

async with httpx.AsyncClient(base_url="http://localhost:8000") as client:
    r = await client.post("/api/wiki/generate", json=payload)
    task_id = r.json()["task_id"]
    print("Task started:", task_id)

```

This creates the Redis entry via `RedisTasks().add_task` as defined in [`api/wiki.py`](https://github.com/quangdungluong/codewiki/blob/main/api/wiki.py) lines 35-44.

### Polling for Task Completion

Monitor the generation progress until completion:

```python
import time
import httpx

def poll(task_id: str):
    with httpx.Client(base_url="http://localhost:8000") as client:
        while True:
            r = client.get(f"/api/wiki/status/{task_id}")
            data = r.json()
            print(f"[{data['status']}] {data['message']}")
            if data["status"] in ("success", "error"):
                break
            time.sleep(2)

poll(task_id)

```

This reads the Redis entry via `RedisTasks().get_task` (see [`api/wiki.py`](https://github.com/quangdungluong/codewiki/blob/main/api/wiki.py) lines 65-71).

### Inspecting Raw Redis State

For debugging, inspect the exact JSON stored for any task:

```python
import redis
import json

r = redis.Redis(host="localhost", port=6379, db=0)
task_key = "<uuid-from-previous-step>"
raw = r.get(task_key)
print(json.loads(raw))

```

This reveals the live task structure containing `status`, `progress`, and partial results.

## Summary

- The **wiki generation pipeline** uses FastAPI `BackgroundTasks` to orchestrate an asynchronous workflow across seven distinct stages, from repository fetching to final cache persistence.
- **Redis task queuing** stores state as JSON keyed by UUID, updated atomically via `RedisTasks.update_task` to prevent race conditions when three concurrent workers generate page content simultaneously.
- Real-time progress tracking is enabled through a polling endpoint at `/api/wiki/status/{task_id}`, which returns the current `status`, `message`, `progress` list, and partial `WikiCacheData` directly from Redis.
- The Gemini model (`gemini-2.5-pro`) drives both wiki structure planning (via XML) and content generation (enforced with specific Markdown formats), processed within [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py).

## Frequently Asked Questions

### Why does CodeWiki use Redis instead of a dedicated message queue like RabbitMQ or Celery?

Redis provides sufficient functionality for this use case without the operational overhead of a heavyweight broker. The `RedisTasks` wrapper leverages atomic `SET`/`GET` operations to maintain a single source of truth for task state, supporting concurrent writes from multiple page-generation workers while keeping the architecture lightweight and simple.

### How does the pipeline handle failures during concurrent page generation?

When a worker encounters an error, it captures the exception details and calls `RedisTasks.update_task` to set the task `status` to `"error"` and populate the `error` field in the JSON payload. Clients polling `/api/wiki/status/{task_id}` receive this error state immediately, allowing them to handle failures gracefully without waiting for the entire batch to complete.

### What specific data is stored in Redis for each generation task?

Each Redis entry contains a JSON object with the task `status` (e.g., `started`, `processing`, `success`), a descriptive `message`, an optional `error` string, a `progress` array listing page IDs currently being processed, and partial `WikiCacheData` containing successfully generated pages. This structure is defined and manipulated throughout [`api/wiki.py`](https://github.com/quangdungluong/codewiki/blob/main/api/wiki.py) and [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py).

### How many concurrent workers does the pipeline support, and where is this configured?

The pipeline defaults to **three concurrent workers** for page content generation. This concurrency level is implemented within the worker pool logic in [`utils/repository_structure.py`](https://github.com/quangdungluong/codewiki/blob/main/utils/repository_structure.py) lines 500-618, where background tasks are spawned to process individual `WikiPage` objects in parallel while respecting the shared Redis state.