# WeKnora Wiki Mode Auto-Generation Pipeline: Architecture and Revision History

> Explore the WeKnora Wiki Mode auto-generation pipeline architecture. Learn how it transforms documents into linked Markdown pages and manages revision history efficiently.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: architecture
- Published: 2026-09-12

---

**WeKnora’s Wiki Mode uses an asynchronous four-stage pipeline orchestrated through Redis to transform ingested documents into inter-linked Markdown pages, storing every modification in a versioned `wiki_page_revisions` table with configurable retention limits.**

The Tencent/WeKnora repository implements a **Wiki Mode** feature that automatically converts knowledge base documents into a structured, browsable wiki. After document ingestion completes, the system triggers a background pipeline that employs multiple LLM prompts to extract entities, build taxonomies, cite source chunks, and deduplicate content before generating the final Markdown pages. Every change—whether generated by the pipeline, an agent tool, or a human editor—is permanently recorded in a dedicated revision history table.

## Four-Stage Wiki Mode Auto-Generation Pipeline

The pipeline is initiated automatically upon document upload completion. It processes documents through five distinct passes (numbered 0 through 4), each implemented as a `TypeWikiIngest` task defined in [`internal/types/task.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/task.go).

### Pass 0: Candidate Slug Extraction

The pipeline begins by scanning the newly ingested document to identify candidate **slugs**—unique entity and concept identifiers—along with a preliminary taxonomy structure. This stage uses the `WikiCandidateSlugPrompt` to extract these identifiers from the raw text.

### Pass 1: Taxonomy Planning

Following extraction, the system organizes candidate slugs into a hierarchical directory structure with a maximum depth of two levels. The `WikiTaxonomyPlanPrompt` ensures the resulting wiki tree maintains logical coherence and proper categorization before any content generation occurs.

### Pass 2: Chunk Citation

For each identified slug, the pipeline walks through document chunks and employs the `WikiChunkCitationPrompt` to map relevant content segments to their source chunk IDs. This enables full traceability, allowing users to verify that generated wiki content is grounded in the original document text. The implementation utilizes prefix caching to optimize LLM token usage during this citation phase.

### Pass 3: Deduplication

Before creating new pages, the system checks whether extracted slugs already exist in the wiki. Using the `WikiDeduplicationPrompt`, it applies the rule that "related ≠ same" to determine if a new slug should merge with an existing page or create a separate entry. This pass returns a mapping of merge decisions that prevents content fragmentation.

### Pass 4: Page Generation and Reduction

The final stage incrementally writes or updates wiki pages using the `WikiPageModifySystemPrompt` and `WikiPageModifyUserPrompt`. For each slug, the LLM generates both a concise summary and the full Markdown body content. This pass specifically enforces constraints against hallucinations and prohibits self-referential links, ensuring generated content remains factual and properly grounded in the source material.

## Asynchronous Task Orchestration

The entire pipeline operates asynchronously through a Redis-backed task queue. When a document finishes ingestion, [`internal/handler/wiki_ingest_batch.go`](https://github.com/Tencent/WeKnora/blob/main/internal/handler/wiki_ingest_batch.go) enqueues a `TypeWikiIngest` task that orchestrates the five passes. Upon completion of a batch, the system records a `wiki_content_changed` activity event.

```go
// Enqueue a Wiki ingest task after a document is uploaded
func enqueueWikiIngest(kbID, docID string) {
    task := tasks.NewTask(
        internal.TaskTypeWikiIngest,
        map[string]string{"kb_id": kbID, "doc_id": docID},
    )
    redisClient.Enqueue(task)
}

```

The task type constants `TypeWikiIngest` and `TypeWikiFinalize` are registered in [`internal/types/task.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/task.go), where the wiki workers poll for new tasks. This decoupled architecture ensures that heavy document processing does not block the API response while maintaining reliability through Redis persistence.

## Revision History Storage Architecture

Every edit to a wiki page—regardless of origin—creates an immutable record in the **`wiki_page_revisions`** table (defined in migration `000075`). The schema captures complete page state snapshots alongside metadata identifying the edit source.

### Database Schema

The revision table stores:

- `page_id` paired with `version` as a composite unique key
- Complete snapshots of `title`, `content`, `summary`, `page_type`, `status`, and `aliases`
- `edit_source` enum distinguishing between `pipeline`, `agent`, `user`, and `revert` origins
- `editor_id` referencing the user or agent responsible
- `edited_at` timestamp for chronological ordering

### Retention Policies

The system enforces two-tiered retention to manage storage growth:

- **Soft limit**: Maintains the 50 most recent versions, specifically preserving pipeline-generated versions and empty-source entries
- **Hard cap**: Absolute maximum of 200 versions per page, preventing unbounded storage expansion regardless of edit frequency

These policies enable the UI to display comprehensive version histories while constraining database size. The [`frontend/src/views/knowledge/wiki/WikiBrowser.vue`](https://github.com/Tencent/WeKnora/blob/main/frontend/src/views/knowledge/wiki/WikiBrowser.vue) component uses this data to present a timeline distinguishing automatic generation events from manual interventions.

```go
// Insert a new revision (called by the pipeline and by manual handlers)
func InsertWikiRevision(ctx context.Context, rev *model.WikiPageRevision) error {
    _, err := db.ExecContext(
        ctx,
        `INSERT INTO wiki_page_revisions
         (page_id, version, title, content, summary, page_type, status,
          edit_source, editor_id, edited_at)
         VALUES ($1,$2,$3,$4,$5,$6,$7,$8,$9,now())`,
        rev.PageID, rev.Version, rev.Title, rev.Content, rev.Summary,
        rev.PageType, rev.Status, rev.EditSource, rev.EditorID,
    )
    return err
}

```

### Querying Revision History

The API exposes endpoints to retrieve version histories, supporting rollback functionality. The following pattern retrieves recent revisions for the wiki browser interface:

```go
// Retrieve the last N revisions for a page (used by the UI)
func ListRecentRevisions(pageID string, limit int) ([]model.WikiPageRevision, error) {
    rows, err := db.Query(`
        SELECT version, title, edit_source, editor_id, edited_at
        FROM wiki_page_revisions
        WHERE page_id = $1
        ORDER BY version DESC
        LIMIT $2`, pageID, limit)
    // … scan rows into a slice …
}

```

## Summary

- **WeKnora’s Wiki Mode** implements a five-pass pipeline (Pass 0-4) that transforms documents into Markdown through progressive LLM prompts for extraction, taxonomy, citation, deduplication, and generation.
- **Task orchestration** relies on Redis-backed asynchronous processing using `TypeWikiIngest` and `TypeWikiFinalize` task types defined in [`internal/types/task.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/task.go).
- **Revision tracking** occurs in the `wiki_page_revisions` table, storing complete snapshots with `edit_source` metadata to distinguish pipeline, agent, user, and revert origins.
- **Storage governance** applies a soft limit of 50 versions and a hard cap of 200 versions per page, ensuring scalable retention of wiki history.
- **End-to-end traceability** connects generated content back to source document chunks through the citation mechanism implemented in Pass 2.

## Frequently Asked Questions

### How many stages does the WeKnora Wiki Mode auto-generation pipeline contain?

The pipeline contains **five sequential passes** (numbered 0 through 4) that progressively transform raw documents into structured wiki pages. Pass 0 extracts candidate slugs, Pass 1 plans the taxonomy, Pass 2 cites source chunks, Pass 3 deduplicates entities, and Pass 4 generates the final Markdown content and summaries.

### What is the maximum number of revisions stored per wiki page?

WeKnora enforces a **hard cap of 200 revisions** per page regardless of edit frequency, supplemented by a **soft limit of 50 recent versions** that prioritizes pipeline-generated snapshots. Once the hard cap is reached, older versions are purged to maintain database performance while preserving the most recent history.

### Can users distinguish between AI-generated and manually edited wiki revisions?

Yes. The `wiki_page_revisions` table includes an **`edit_source`** column that explicitly tags each version as `pipeline` (auto-generated), `agent` (tool-assisted), `user` (manual edit), or `revert` (rollback operation). The frontend displays these sources in the version history browser, allowing users to identify the origin of each change.

### How does the pipeline prevent duplicate wiki pages for the same entity?

During **Pass 3 (Deduplication)**, the pipeline invokes the `WikiDeduplicationPrompt` to evaluate whether a newly extracted slug represents the same concept as an existing page. By applying the principle that "related ≠ same," the system either maps the new content to an existing page or creates a separate entry, preventing fragmentation while allowing distinct but related concepts to coexist.