Wiki Generation Pipeline and Redis Task Queuing in CodeWiki: A Technical Deep Dive
The wiki generation pipeline is an asynchronous, multi-stage workflow that transforms code repositories into structured technical wikis, using Redis as a lightweight state store to track progress and coordinate concurrent workers without requiring a dedicated message broker.
The quangdungluong/codewiki project implements a documentation engine that converts source code into AI-generated wikis through a sophisticated wiki generation pipeline. This system leverages FastAPI endpoints and Redis task queuing to orchestrate repository analysis, LLM-powered content creation, and real-time status delivery.
Seven Stages of the Wiki Generation Pipeline
Stage 1: Task Initialization
When a client POSTs a WikiTaskRequest to /api/wiki/generate, the system generates a UUID and writes an initial task entry to Redis with status started. FastAPI’s BackgroundTasks then hands the request to the processing engine. This logic resides in [api/wiki.py](https://github.com/quangdungluong/codewiki/blob/master/api/wiki.py) lines 33-48.
Stage 2: Repository Structure Fetching
The RepositoryStructureFetcher contacts the appropriate source—whether a local API or the GitHub API—to build a complete file tree and README snapshot. This phase executes within [utils/repository_structure.py](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 68-115.
Stage 3: Wiki Structure Determination
The fetched repository data is fed to the Gemini model (gemini-2.5-pro), which returns an XML description of pages, sections, and their hierarchical relationships. The system parses this XML into internal Pydantic models: WikiStructure, WikiPage, and WikiSection. See [utils/repository_structure.py](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 180-388.
Stage 4: Concurrent Page Content Generation
For every page defined in the structure, background workers invoke the Gemini model with detailed prompts enforcing specific Markdown formats like <details> blocks, Mermaid diagrams, and tables. By default, three workers run concurrently to generate content in parallel. This concurrency logic appears in [utils/repository_structure.py](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 500-618.
Stage 5: Real-Time Status Persistence
After each significant step—fetching, structure building, or page completion—the pipeline calls RedisTasks.update_task to mutate the JSON stored under the task UUID. This JSON contains status, message, optional error details, partial WikiCacheData, and a progress list tracking currently processing page IDs. The update methods are located in [api/wiki.py](https://github.com/quangdungluong/codewiki/blob/master/api/wiki.py) lines 15-31, 76-84, and 122-130.
Stage 6: Client Polling Interface
Clients monitor progress by querying /api/wiki/status/{task_id}, which reads the latest task entry from Redis via RedisTasks.get_task and returns the current JSON payload. This endpoint is implemented in [api/wiki.py](https://github.com/quangdungluong/codewiki/blob/master/api/wiki.py) lines 65-71.
Stage 7: Cache Persistence
Upon successful completion, the assembled WikiStructure and generated page contents are POSTed to the cache service at /api/wiki_cache. This final persistence step is handled in [utils/repository_structure.py](https://github.com/quangdungluong/codewiki/blob/master/utils/repository_structure.py) lines 438-447.
Redis Task State Management
The system uses Redis not as a traditional message queue, but as an atomic state store. The RedisTasks class in [utils/redis_tasks.py](https://github.com/quangdungluong/codewiki/blob/master/utils/redis_tasks.py) provides a thin wrapper around redis-py, implementing add_task, get_task, update_task, and delete_task operations.
Each task is stored as a JSON string keyed by its UUID. Unlike systems using separate queue data structures, the task record itself acts as the single source of truth for status and progress. Because Redis supports atomic SET and GET operations, multiple concurrent workers can safely update task state without race conditions. The background workers write updates (e.g., status="processing", progress=[...]) while the API endpoints read the current values.
Core Implementation Files
| File | Role |
|---|---|
api/wiki.py |
FastAPI router handling task creation at /api/wiki/generate, status polling at /api/wiki/status/{task_id}, and background task orchestration. |
utils/redis_tasks.py |
Redis wrapper managing atomic task state operations (add_task, get_task, update_task, delete_task). |
utils/repository_structure.py |
Core engine fetching repository data, invoking the LLM for structure generation, managing concurrent page workers, and updating Redis state throughout execution. |
utils/document_pipeline.py |
Provides RecursiveDocumentReader and DocumentTransformer for source file processing used during content generation. |
utils/models.py |
Pydantic models (WikiPage, WikiSection, WikiStructure, WikiCacheData) for data validation and serialization across the pipeline. |
Code Examples: Interacting with the Pipeline
Triggering a New Wiki Generation
Use httpx to submit a repository for processing:
import httpx
import uuid
payload = {
"repo_info": {"type": "web"},
"repo_url": "https://github.com/quangdungluong/codewiki",
"owner": "quangdungluong",
"repo": "codewiki",
"token": None,
}
async with httpx.AsyncClient(base_url="http://localhost:8000") as client:
r = await client.post("/api/wiki/generate", json=payload)
task_id = r.json()["task_id"]
print("Task started:", task_id)
This creates the Redis entry via RedisTasks().add_task as defined in api/wiki.py lines 35-44.
Polling for Task Completion
Monitor the generation progress until completion:
import time
import httpx
def poll(task_id: str):
with httpx.Client(base_url="http://localhost:8000") as client:
while True:
r = client.get(f"/api/wiki/status/{task_id}")
data = r.json()
print(f"[{data['status']}] {data['message']}")
if data["status"] in ("success", "error"):
break
time.sleep(2)
poll(task_id)
This reads the Redis entry via RedisTasks().get_task (see api/wiki.py lines 65-71).
Inspecting Raw Redis State
For debugging, inspect the exact JSON stored for any task:
import redis
import json
r = redis.Redis(host="localhost", port=6379, db=0)
task_key = "<uuid-from-previous-step>"
raw = r.get(task_key)
print(json.loads(raw))
This reveals the live task structure containing status, progress, and partial results.
Summary
- The wiki generation pipeline uses FastAPI
BackgroundTasksto orchestrate an asynchronous workflow across seven distinct stages, from repository fetching to final cache persistence. - Redis task queuing stores state as JSON keyed by UUID, updated atomically via
RedisTasks.update_taskto prevent race conditions when three concurrent workers generate page content simultaneously. - Real-time progress tracking is enabled through a polling endpoint at
/api/wiki/status/{task_id}, which returns the currentstatus,message,progresslist, and partialWikiCacheDatadirectly from Redis. - The Gemini model (
gemini-2.5-pro) drives both wiki structure planning (via XML) and content generation (enforced with specific Markdown formats), processed withinutils/repository_structure.py.
Frequently Asked Questions
Why does CodeWiki use Redis instead of a dedicated message queue like RabbitMQ or Celery?
Redis provides sufficient functionality for this use case without the operational overhead of a heavyweight broker. The RedisTasks wrapper leverages atomic SET/GET operations to maintain a single source of truth for task state, supporting concurrent writes from multiple page-generation workers while keeping the architecture lightweight and simple.
How does the pipeline handle failures during concurrent page generation?
When a worker encounters an error, it captures the exception details and calls RedisTasks.update_task to set the task status to "error" and populate the error field in the JSON payload. Clients polling /api/wiki/status/{task_id} receive this error state immediately, allowing them to handle failures gracefully without waiting for the entire batch to complete.
What specific data is stored in Redis for each generation task?
Each Redis entry contains a JSON object with the task status (e.g., started, processing, success), a descriptive message, an optional error string, a progress array listing page IDs currently being processed, and partial WikiCacheData containing successfully generated pages. This structure is defined and manipulated throughout api/wiki.py and utils/repository_structure.py.
How many concurrent workers does the pipeline support, and where is this configured?
The pipeline defaults to three concurrent workers for page content generation. This concurrency level is implemented within the worker pool logic in utils/repository_structure.py lines 500-618, where background tasks are spawned to process individual WikiPage objects in parallel while respecting the shared Redis state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →