How to Generate a Wiki from Code Communities with Wikilinks in Code-Review-Graph
The code-review-graph repository converts detected code communities into a navigable Markdown wiki by slugifying community names into filenames and linking them with standard Markdown wikilinks in code_review_graph/wiki.py.
This open-source tool analyzes codebases as graphs, detects logical communities, and exports them as a browsable documentation set. The wiki generation pipeline resides in code_review_graph/wiki.py, which consumes community metadata and produces interlinked Markdown files without requiring a custom parser.
The Wiki Generation Pipeline
The orchestration happens inside the generate_wiki function. It transforms abstract community data into concrete documentation through a deterministic, collision-safe process.
Collecting Communities from the Graph
The pipeline begins by fetching detected communities. The generate_wiki function calls get_communities(store), imported from communities.py, which returns a list of dictionaries containing community names, sizes, and member files.
from code_review_graph.wiki import generate_wiki
from code_review_graph.graph import GraphStore
store = GraphStore(path="path/to/graph.db")
wiki_dir = "/tmp/my-project-wiki"
result = generate_wiki(store, wiki_dir, force=False)
This decoupled architecture means wiki.py depends only on the public get_communities API, keeping the documentation generator independent of the underlying graph detection algorithms.
Slugifying Community Names
Raw community names often contain spaces, punctuation, or Unicode characters unsafe for filenames. The _slugify helper (defined at line 24 in wiki.py) normalizes Unicode, replaces non-alphanumeric characters with hyphens, truncates to 80 characters, and guarantees a non-empty fallback for edge cases like emoji-only names.
Handling Name Collisions
When two distinct communities resolve to identical slugs, the generator appends numeric suffixes (-2, -3, etc.) until each filename is unique. This collision handling loop (lines 213-219) prevents file overwrites and ensures stable URLs across regeneration runs, which is critical for incremental updates in CI pipelines.
Writing Community Pages with Wikilinks
For each community, the system writes a dedicated <slug>.md file containing:
- A header with the community name
- A description (if metadata includes one)
- A table listing member files
- Wikilinks referencing dependent communities using standard Markdown syntax:
[other-community.md](other-community.md)
This approach requires no special rendering engine; GitHub, VS Code, and MkDocs handle the links natively.
Building the Index Page
After emitting all community pages, generate_wiki creates index.md at the wiki root. The index features a Markdown table with three columns—Community, Size, and Page—linking each entry to its slugified filename (lines 255-256). This provides immediate navigation into the community structure.
Security and Incremental Updates
Path-Traversal Protection
The get_wiki_page function (lines 280-303) implements strict read-only retrieval. It resolves a requested page name to its slug, then validates that the resolved path remains inside the wiki directory boundary, defending against path-traversal attacks.
Incremental Generation
Instead of rewriting every file on each invocation, the generator checks timestamps and file hashes via os.path.getmtime. When force=False, it skips unchanged files, returning a summary dictionary indicating counts for generated, updated, and unchanged pages. This optimization makes the tool suitable for large repositories and frequent CI jobs.
Practical Implementation Examples
Generate a Complete Wiki
from code_review_graph.wiki import generate_wiki
from code_review_graph.graph import GraphStore
store = GraphStore(path="path/to/graph.db")
wiki_dir = "/tmp/my-project-wiki"
result = generate_wiki(store, wiki_dir, force=False)
print(result)
# → {'pages_generated': 12, 'pages_updated': 0, 'pages_unchanged': 12}
This creates index.md plus individual <slug>.md files for each detected community inside wiki_dir.
Retrieve a Specific Page Safely
from code_review_graph.wiki import get_wiki_page
content = get_wiki_page("/tmp/my-project-wiki", "Data Processing")
if content:
print(content)
else:
print("Page not found")
The function safely resolves "Data Processing" to data-processing.md while validating the path does not escape the wiki root.
CLI Integration
The functionality is exposed via code_review_graph/tools/docs.py for command-line usage:
crg generate-wiki --store path/to/graph.db --out ./wiki
This wrapper calls generate_wiki and prints a concise summary suitable for automated documentation pipelines.
Summary
- Wiki generation is orchestrated in
code_review_graph/wiki.pyvia thegenerate_wikifunction, which consumes output fromcommunities.py. - Slugification converts community names to safe filenames using
_slugify, with collision handling to prevent overwrites. - Wikilinks use standard Markdown syntax
[filename.md](filename.md), ensuring compatibility with existing documentation tools. - Security is enforced in
get_wiki_pagethrough strict path validation against directory traversal attacks. - Incremental updates skip unchanged files based on timestamps, optimizing CI performance.
Frequently Asked Questions
How does code-review-graph handle duplicate community names during wiki generation?
When the slugification process produces identical filenames for different communities, the collision handler appends incremental numeric suffixes (-2, -3, etc.) until each filename is unique. This logic at lines 213-219 of wiki.py ensures no data is overwritten while maintaining deterministic naming.
What file format does the wiki generator use for inter-page links?
The generator uses standard Markdown link syntax: [target-slug.md](target-slug.md). These wikilinks require no custom parser and render correctly in GitHub, VS Code, MkDocs, and other standard Markdown viewers.
Can the wiki be regenerated incrementally without rewriting every file?
Yes. The generate_wiki function accepts a force parameter. When set to False, it compares file modification times and hashes, only updating files that have changed. This returns a summary dictionary indicating how many pages were generated, updated, or left unchanged.
How does the system prevent directory traversal when retrieving wiki pages?
The get_wiki_page function strictly checks that the resolved file path remains within the wiki directory boundary before reading. If the requested page name resolves to a path outside the root (e.g., containing ../), the function returns None to block the attack.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →