WeKnora Wiki Mode Auto-Generation Pipeline: Architecture and Revision History
WeKnora’s Wiki Mode uses an asynchronous four-stage pipeline orchestrated through Redis to transform ingested documents into inter-linked Markdown pages, storing every modification in a versioned wiki_page_revisions table with configurable retention limits.
The Tencent/WeKnora repository implements a Wiki Mode feature that automatically converts knowledge base documents into a structured, browsable wiki. After document ingestion completes, the system triggers a background pipeline that employs multiple LLM prompts to extract entities, build taxonomies, cite source chunks, and deduplicate content before generating the final Markdown pages. Every change—whether generated by the pipeline, an agent tool, or a human editor—is permanently recorded in a dedicated revision history table.
Four-Stage Wiki Mode Auto-Generation Pipeline
The pipeline is initiated automatically upon document upload completion. It processes documents through five distinct passes (numbered 0 through 4), each implemented as a TypeWikiIngest task defined in internal/types/task.go.
Pass 0: Candidate Slug Extraction
The pipeline begins by scanning the newly ingested document to identify candidate slugs—unique entity and concept identifiers—along with a preliminary taxonomy structure. This stage uses the WikiCandidateSlugPrompt to extract these identifiers from the raw text.
Pass 1: Taxonomy Planning
Following extraction, the system organizes candidate slugs into a hierarchical directory structure with a maximum depth of two levels. The WikiTaxonomyPlanPrompt ensures the resulting wiki tree maintains logical coherence and proper categorization before any content generation occurs.
Pass 2: Chunk Citation
For each identified slug, the pipeline walks through document chunks and employs the WikiChunkCitationPrompt to map relevant content segments to their source chunk IDs. This enables full traceability, allowing users to verify that generated wiki content is grounded in the original document text. The implementation utilizes prefix caching to optimize LLM token usage during this citation phase.
Pass 3: Deduplication
Before creating new pages, the system checks whether extracted slugs already exist in the wiki. Using the WikiDeduplicationPrompt, it applies the rule that "related ≠ same" to determine if a new slug should merge with an existing page or create a separate entry. This pass returns a mapping of merge decisions that prevents content fragmentation.
Pass 4: Page Generation and Reduction
The final stage incrementally writes or updates wiki pages using the WikiPageModifySystemPrompt and WikiPageModifyUserPrompt. For each slug, the LLM generates both a concise summary and the full Markdown body content. This pass specifically enforces constraints against hallucinations and prohibits self-referential links, ensuring generated content remains factual and properly grounded in the source material.
Asynchronous Task Orchestration
The entire pipeline operates asynchronously through a Redis-backed task queue. When a document finishes ingestion, internal/handler/wiki_ingest_batch.go enqueues a TypeWikiIngest task that orchestrates the five passes. Upon completion of a batch, the system records a wiki_content_changed activity event.
// Enqueue a Wiki ingest task after a document is uploaded
func enqueueWikiIngest(kbID, docID string) {
task := tasks.NewTask(
internal.TaskTypeWikiIngest,
map[string]string{"kb_id": kbID, "doc_id": docID},
)
redisClient.Enqueue(task)
}
The task type constants TypeWikiIngest and TypeWikiFinalize are registered in internal/types/task.go, where the wiki workers poll for new tasks. This decoupled architecture ensures that heavy document processing does not block the API response while maintaining reliability through Redis persistence.
Revision History Storage Architecture
Every edit to a wiki page—regardless of origin—creates an immutable record in the wiki_page_revisions table (defined in migration 000075). The schema captures complete page state snapshots alongside metadata identifying the edit source.
Database Schema
The revision table stores:
page_idpaired withversionas a composite unique key- Complete snapshots of
title,content,summary,page_type,status, andaliases edit_sourceenum distinguishing betweenpipeline,agent,user, andrevertoriginseditor_idreferencing the user or agent responsibleedited_attimestamp for chronological ordering
Retention Policies
The system enforces two-tiered retention to manage storage growth:
- Soft limit: Maintains the 50 most recent versions, specifically preserving pipeline-generated versions and empty-source entries
- Hard cap: Absolute maximum of 200 versions per page, preventing unbounded storage expansion regardless of edit frequency
These policies enable the UI to display comprehensive version histories while constraining database size. The frontend/src/views/knowledge/wiki/WikiBrowser.vue component uses this data to present a timeline distinguishing automatic generation events from manual interventions.
// Insert a new revision (called by the pipeline and by manual handlers)
func InsertWikiRevision(ctx context.Context, rev *model.WikiPageRevision) error {
_, err := db.ExecContext(
ctx,
`INSERT INTO wiki_page_revisions
(page_id, version, title, content, summary, page_type, status,
edit_source, editor_id, edited_at)
VALUES ($1,$2,$3,$4,$5,$6,$7,$8,$9,now())`,
rev.PageID, rev.Version, rev.Title, rev.Content, rev.Summary,
rev.PageType, rev.Status, rev.EditSource, rev.EditorID,
)
return err
}
Querying Revision History
The API exposes endpoints to retrieve version histories, supporting rollback functionality. The following pattern retrieves recent revisions for the wiki browser interface:
// Retrieve the last N revisions for a page (used by the UI)
func ListRecentRevisions(pageID string, limit int) ([]model.WikiPageRevision, error) {
rows, err := db.Query(`
SELECT version, title, edit_source, editor_id, edited_at
FROM wiki_page_revisions
WHERE page_id = $1
ORDER BY version DESC
LIMIT $2`, pageID, limit)
// … scan rows into a slice …
}
Summary
- WeKnora’s Wiki Mode implements a five-pass pipeline (Pass 0-4) that transforms documents into Markdown through progressive LLM prompts for extraction, taxonomy, citation, deduplication, and generation.
- Task orchestration relies on Redis-backed asynchronous processing using
TypeWikiIngestandTypeWikiFinalizetask types defined ininternal/types/task.go. - Revision tracking occurs in the
wiki_page_revisionstable, storing complete snapshots withedit_sourcemetadata to distinguish pipeline, agent, user, and revert origins. - Storage governance applies a soft limit of 50 versions and a hard cap of 200 versions per page, ensuring scalable retention of wiki history.
- End-to-end traceability connects generated content back to source document chunks through the citation mechanism implemented in Pass 2.
Frequently Asked Questions
How many stages does the WeKnora Wiki Mode auto-generation pipeline contain?
The pipeline contains five sequential passes (numbered 0 through 4) that progressively transform raw documents into structured wiki pages. Pass 0 extracts candidate slugs, Pass 1 plans the taxonomy, Pass 2 cites source chunks, Pass 3 deduplicates entities, and Pass 4 generates the final Markdown content and summaries.
What is the maximum number of revisions stored per wiki page?
WeKnora enforces a hard cap of 200 revisions per page regardless of edit frequency, supplemented by a soft limit of 50 recent versions that prioritizes pipeline-generated snapshots. Once the hard cap is reached, older versions are purged to maintain database performance while preserving the most recent history.
Can users distinguish between AI-generated and manually edited wiki revisions?
Yes. The wiki_page_revisions table includes an edit_source column that explicitly tags each version as pipeline (auto-generated), agent (tool-assisted), user (manual edit), or revert (rollback operation). The frontend displays these sources in the version history browser, allowing users to identify the origin of each change.
How does the pipeline prevent duplicate wiki pages for the same entity?
During Pass 3 (Deduplication), the pipeline invokes the WikiDeduplicationPrompt to evaluate whether a newly extracted slug represents the same concept as an existing page. By applying the principle that "related ≠ same," the system either maps the new content to an existing page or creates a separate entry, preventing fragmentation while allowing distinct but related concepts to coexist.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →