Understanding the Three-Layer Architecture of LLM Wiki: Raw Sources, Wiki, and Schema

LLM Wiki implements a three-layer architecture separating immutable raw sources, LLM-generated wiki pages, and declarative schema rules to create a traceable, evolvable knowledge base.

The nashsu/llm_wiki repository structures personal knowledge management around a clean separation of concerns inspired by Andrej Karpathy’s design patterns. This three-layer architecture ensures that your original documents remain untouched while the LLM generates enriched, interconnected content governed by explicit rules.

The Three Layers Explained

Raw Sources Layer

The Raw Sources layer stores immutable input documents that serve as the single source of truth. Located at raw/sources/, this directory contains PDFs, DOCX files, Markdown notes, images, and web clips.

These files are never edited by the LLM. When you add a new research paper or set of notes, you place it directly in raw/sources/ and trigger the ingest pipeline. The system treats these documents as immutable ground truth, ensuring that any information derived from them can be traced back to the original file.

Wiki Layer

The Wiki layer contains LLM-generated knowledge pages that summarize, enrich, and interlink the raw material. Files live in wiki/ and include entities (people, organizations), concepts (abstract ideas), source summaries, index.md (the central catalog), and log.md (the chronological operation record).

Each wiki page contains YAML front matter linking back to its raw sources, typically via a sources: [] array. This creates a bidirectional relationship: the wiki page references its origins in raw/sources/, while the ingest pipeline updates wiki/index.md to track all generated pages. This design enables incremental updates—only changed source files trigger re-ingestion, leaving the rest of the wiki untouched.

Schema Layer

The Schema layer defines the rules, configuration, and purpose that govern LLM behavior. This layer resides in two critical files at the repository root:

  • purpose.md declares project goals and research scope, guiding the LLM’s context during both ingest and query phases.
  • schema.md defines page-type rules, validation requirements, and allowed link patterns.

This declarative approach provides human-in-the-loop control. Rather than modifying Python code, you steer the system’s behavior by editing Markdown files. For example, schema.md might specify that an entity page requires title, description, and sources fields, while permitting links to concept or entity pages.

Architecture in Practice

Directory Structure

A minimal LLM Wiki project demonstrates the separation of concerns:

my-wiki/
├── purpose.md            # 📄 Schema layer: project goals

├── schema.md            # 📄 Schema layer: page-type definitions

├── raw/
│   └── sources/         # 📁 Raw Sources layer: immutable inputs

│       ├── paper.pdf
│       └── notes.md
└── wiki/
    ├── index.md         # 📄 Wiki layer: central catalog

    ├── log.md           # 📄 Wiki layer: operation chronology

    ├── entities/
    │   └── alice.md
    └── concepts/
        └── quantum.md

To add content, place a file in raw/sources/ and run the ingest process. The LLM generates appropriate wiki pages under wiki/ and updates wiki/index.md automatically.

Schema Configuration

The schema.md file enforces structure through declarative rules:


# schema.md – page-type definitions

pageTypes:
  entity:
    requiredFields: [title, description, sources]
    allowedLinks: [concept, entity]
  concept:
    requiredFields: [title, overview, sources]
    allowedLinks: [entity, concept]

These rules validate generated pages before writing, ensuring that every entity page contains the required metadata and only links to permitted page types.

Runtime Query Flow

When interacting with the local HTTP API, all three layers collaborate:

curl -X POST http://127.0.0.1:19828/api/v1/projects/1/chat \
  -H "Content-Type: application/json" \
  -d '{
        "messages": [{"role":"user","content":"Explain quantum entanglement"}],
        "stream": false
      }'

This request executes in two phases:

  1. Phase 1 – Retrieval: The system searches the wiki/ layer for relevant concept pages.
  2. Phase 2 – LLM Generation: The LLM reads the retrieved pages and the schema.md and purpose.md files to generate a concise answer citing the sources.

Key Files and Their Roles

File Layer Function
raw/sources/* Raw Sources Immutable ground truth documents stored under version control.
wiki/index.md Wiki Central catalog enabling fast lookup and graph construction.
wiki/log.md Wiki Chronological audit trail of all ingest operations.
purpose.md Schema Declares research scope and project goals referenced during generation.
schema.md Schema Defines validation rules and page-type constraints.
README.md Documentation Documents the three-layer architecture (lines 71-74).

Summary

  • Immutable foundation: The raw/sources/ directory preserves original documents as the single source of truth.
  • Generated knowledge: The wiki/ layer contains LLM-created, interlinked pages with YAML front matter tracing back to raw sources.
  • Declarative control: schema.md and purpose.md govern LLM behavior without requiring code changes.
  • Traceability: Every generated page links to its source files, enabling full lineage tracking.
  • Incremental updates: Only modified raw files trigger re-ingestion, preserving existing wiki content.

Frequently Asked Questions

What is the three-layer architecture of LLM Wiki?

The three-layer architecture consists of Raw Sources (immutable input files in raw/sources/), the Wiki (LLM-generated pages in wiki/ with metadata linking to sources), and the Schema (governance files schema.md and purpose.md). This separation ensures traceability, incremental updates, and human control over LLM behavior.

How does LLM Wiki maintain traceability between sources and generated content?

Every page in the wiki/ directory includes YAML front matter with a sources: [] array referencing specific files in raw/sources/. This metadata creates a bidirectional link: the wiki page cites its origins, while the ingest pipeline updates wiki/index.md to track all pages derived from each source file.

Can I modify the behavior of LLM Wiki without changing code?

Yes. The Schema layer uses declarative Markdown files. By editing schema.md (to change page-type rules) or purpose.md (to adjust project goals), you steer LLM generation during both ingest and query phases. This human-in-the-loop approach requires no Python modifications.

How does the architecture support incremental updates?

Because raw sources are immutable and wiki pages are generated, the system compares checksums or timestamps of files in raw/sources/ against the wiki/log.md audit trail. Only new or modified sources trigger re-ingestion, leaving existing wiki pages untouched and preserving manual edits or enrichments you may have added.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →