How Wiki and CodeGraph Indexing and Retrieval Work in TencentDB Agent Memory
TencentDB Agent Memory uses a unified serial queue architecture to index Wiki markdown pages and CodeGraph git repositories, storing both as tenant-isolated assets with background workers handling LLM ingestion or graph construction while exposing safe, transactional retrieval APIs.
TencentDB Agent Memory manages two core knowledge asset types: Wiki for structured documentation and CodeGraph for repository code relationships. Both follow identical lifecycle patterns implemented in MemoryKnowledge/src/store/wiki-service.ts and MemoryKnowledge/src/store/code-graph-service.ts respectively, using per-asset build queues to ensure race-free background processing.
Architectural Overview
Both asset types share a five-stage pipeline:
| Stage | Wiki Implementation | CodeGraph Implementation |
|---|---|---|
| Create | WikiService.create() writes WikiRow + directory skeleton |
CodeGraphService.create() writes CodeGraphRow + enqueues build |
| Queue | WikiService.enqueueBuild adds to per-wiki BuildQueue |
CodeGraphService.enqueueBuild uses same queue mechanism |
| Worker | WikiWorker processes raw files → LLM summary → markdown pages |
CodeGraphWorker clones repo, runs graph engine, writes index |
| Status | draft → pending → processing → ready | failed |
pending → processing → ready | failed |
| Retrieval | pageLs(), pageRead(), rawLs(), rawRead*() |
get() returns metadata; graph accessed via worker or disk |
Tenant isolation is enforced through directory hierarchy: <dataRoot>/<service_id>/<team_id>/<asset_id>/.
Wiki Indexing Pipeline
Metadata Creation and Directory Setup
The indexing process begins with WikiService.create() at lines 41-55 of wiki-service.ts. This method:
- Persists a
WikiRowto the SQLite-backedIKnowledgeStore - Creates the filesystem structure:
raw/for source files,wiki/for generated pages
const { row } = wiki.create({
service_id: "svc-001",
team_id: "team-42",
name: "Redis Guide",
});
Serial Build Queue Processing
Each wiki maintains its own BuildQueue to prevent concurrent modifications. WikiService.enqueueBuild (lines 110-112) serializes jobs:
- Status transitions to
pending - The injected
WikiWorkerreceives aWikiBuildContextcontaining:dir: asset directory pathsetInternalStatus: callback for granular progress (scanning, ingesting)ingestRunId: tracking identifier for the LLM pipeline
LLM Ingestion and Search Index Population
Inside the worker, raw files undergo LLM processing to generate summaries and structured markdown. The worker writes outputs via pageWrite* methods and populates the source database through:
initIndexDb: initializes SQLite schema for raw file metadataupsertSource: inserts/updates file records feeding full-text search
Status updates flow through store.updateWikiStatus as shown in the runBuild implementation (lines 14-30).
Wiki Retrieval APIs
Reading Raw Source Files
rawLs() and rawRead*() (lines 102-108) query the source database directly:
const sources = wiki.rawLs("svc-001", "team-42", "wiki-123");
// Returns: array of source file metadata from SQLite source DB
These methods use listSources and getReadDb from MemoryKnowledge/src/engines/wiki/index-db.ts.
Reading Generated Pages
pageLs() (lines 152-162) scans the wiki/ directory, parsing frontmatter with parseFrontmatterMin:
const pages = wiki.pageLs("svc-001", "team-42", "wiki-123");
for (const p of pages) {
const markdown = wiki.pageRead("svc-001", "team-42", "wiki-123", p.id);
}
pageRead() (lines 166-174) uses resolvePageRef for safe path resolution, preventing directory traversal attacks.
CodeGraph Indexing Pipeline
Repository Cloning and Graph Construction
CodeGraphService.create() (lines 32-38) immediately enqueues a build. The CodeGraphWorker receives a CodeGraphBuildContext and:
- Clones the git repository to the checkout directory
- Executes the graph builder engine
- Returns statistics:
commitHash,files,nodes,edges
Status progression mirrors Wiki: internal_status tracks cloning → indexing → ready.
Build Completion and Summary Generation
Upon success, onBuildComplete (lines 65-71) generates a textual summary without LLM overhead:
{
commitHash: "a1b2c3",
stats: { files: 230, nodes: 5400, edges: 12000 }
}
This summary is persisted via updateCodeGraphStatus.
CodeGraph Retrieval and Lifecycle Management
Metadata and Graph Access
codeGraphService.get()/getById()return theCodeGraphRowmetadata- The actual graph resides on disk in the asset directory
- Clients access it through the injected worker or direct file reads
Forced Resynchronization
The sync() method checks current status, marks the asset pending, and enqueues a fresh build—useful for tracking upstream repository changes.
Safe Deletion
cleanupResources (lines 24-31) implements cancellation-aware cleanup:
- Sets
cancelledflag to stop active workers - Releases in-memory graph instance via
releaseInstance - Removes database rows
- Recursively deletes asset directory
Practical Implementation Examples
Complete Wiki Creation and Ingestion
import { WikiService, type WikiWorker } from "./store/wiki-service.ts";
const worker: WikiWorker = async (ctx) => {
// LLM processing: ctx.dir/raw/* → ctx.dir/wiki/*.md
const files = await glob(`${ctx.dir}/raw/**/*`);
for (const f of files) {
const summary = await llm.summarize(f);
await Deno.writeTextFile(
`${ctx.dir}/wiki/${basename(f)}.md`,
`---\ntitle: ${summary.title}\n---\n${summary.body}`
);
}
ctx.setInternalStatus("finalizing");
return { pageCount: files.length };
};
const wiki = new WikiService({ store, dataRoot: "/data", worker });
const { row } = wiki.create({
service_id: "svc-001",
team_id: "team-42",
name: "PostgreSQL Troubleshooting"
});
wiki.ingest(row.service_id, row.team_id, row.wiki_id);
CodeGraph with Custom Worker
import { CodeGraphService, type CodeGraphWorker } from "./store/code-graph-service.ts";
const cgWorker: CodeGraphWorker = async (ctx) => {
const repo = await git.clone(ctx.repoUrl, ctx.dir, { branch: ctx.branch });
const graph = await codeGraphEngine.build(ctx.dir, {
include: ["*.ts", "*.js"],
exclude: ["node_modules/**"]
});
return {
commitHash: repo.head,
stats: graph.stats()
};
};
const cg = new CodeGraphService({ store, dataRoot: "/data", worker: cgWorker });
const { row } = cg.create({
service_id: "svc-001",
team_id: "team-42",
repo_url: "https://github.com/TencentCloud/tmysqlparser.git",
branch: "main"
});
Key Source Files
| Component | Path |
|---|---|
| Wiki service orchestration | MemoryKnowledge/src/store/wiki-service.ts |
| Wiki source database & helpers | MemoryKnowledge/src/engines/wiki/index-db.ts |
| LLM ingestion pipeline | MemoryKnowledge/src/engines/wiki/ingest-v2/* |
| CodeGraph service | MemoryKnowledge/src/store/code-graph-service.ts |
| Serial build queue | MemoryKnowledge/src/store/build-queue.ts |
| Type definitions | MemoryKnowledge/src/store/types.ts |
Summary
- Unified architecture: Both Wiki and CodeGraph use
BuildQueuefor serialized background processing, preventing race conditions - Tenant isolation: Directory hierarchy
<dataRoot>/<service_id>/<team_id>/<asset_id>enforces multi-tenancy - Status-driven lifecycle: Explicit state machines (
pending → processing → ready) enable reliable progress tracking and cancellation - Safe retrieval: Path traversal protection via
resolvePageRefandresolveRawPath; transactional reads returnnullfor missing assets - Idempotent operations: Creation, deletion, and sync operations can be retried safely without data corruption
Frequently Asked Questions
What database does TencentDB Agent Memory use for search indexing?
The system uses SQLite for metadata and source file indexing. The index-db.ts module manages schema initialization (initIndexDb) and CRUD operations (upsertSource, listSources). Full-text search capabilities are built on top of this SQLite layer rather than requiring external search engines.
How does the build queue prevent concurrent modifications to the same asset?
Each wiki and code-graph maintains its own BuildQueue instance. When enqueueBuild is called, jobs are serialized per-asset—subsequent calls for the same wiki_id or code_graph_id wait for prior jobs to complete. This guarantees that background workers never execute concurrently against the same filesystem directory.
Can I retrieve a CodeGraph's graph data directly without going through the worker?
Yes. While get() returns only metadata, the actual graph files are stored on disk under the asset directory (<dataRoot>/<service_id>/<team_id>/<code_graph_id>/). Client code can read these files directly. The injected CodeGraphWorker provides a convenience abstraction but is not required for read access.
What happens if a wiki ingestion fails midway through processing?
The status transitions to failed and the internal_status captures the specific failure point (e.g., "ingesting"). The partial state remains on disk—the wiki/ directory may contain incomplete pages, and the source DB may have partial records. Calling ingest() again resets to pending and re-runs the full pipeline, overwriting previous results idempotently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →