What Is the Multi-Pass Indexing Pipeline in CBM?
The multi-pass indexing pipeline in CBM is a sequential orchestration engine that processes a codebase through discrete analysis stages, allowing each pass to write artifacts that subsequent passes can read and enrich.
The DeusData/codebase-memory-mcp repository implements this architecture to transform raw source files into a queryable knowledge graph. By treating analysis as a series of discrete passes rather than a monolithic procedure, the multi-pass indexing pipeline enables incremental enrichment of codebase metadata while maintaining clear separation of concerns.
Core Architecture of the Multi-Pass Pipeline
At the heart of the system sits the pipeline abstraction defined in [src/pipeline/pipeline.h](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.h). The structure stores the working directory, a dynamic array of registered passes, and their associated contexts.
Pipeline State Management
The constructor pipeline_new() allocates a pipeline_t instance and captures the target repository path:
pipeline_t *pl = pipeline_new("/absolute/path/to/repo");
Internally, this initializes pass_count to zero and prepares empty arrays for pass functions and contexts. The destructor pipeline_free() releases these allocations when analysis completes.
Pass Registration Interface
Contributors add analysis steps via pipeline_register_pass(), which accepts a function pointer conforming to the signature void (*)(pipeline_t *, void *) and an optional context pointer:
void pipeline_register_pass(pipeline_t *pl,
void (*pass)(pipeline_t *, void *),
void *ctx);
The implementation appends the function to the internal passes array and stores the context in a parallel contexts array, preserving insertion order for deterministic execution.
Sequential Execution of Analysis Passes
The pipeline_run() function drives the multi-pass indexing pipeline by iterating through the registered passes in order:
int pipeline_run(pipeline_t *pl) {
for (size_t i = 0; i < pl->pass_count; ++i) {
void (*fn)(pipeline_t *, void *) = pl->passes[i].func;
void *ctx = pl->contexts[i];
fn(pl, ctx);
}
return 0;
}
Each pass receives the global pipeline state, enabling it to read artifacts produced by earlier passes and append new ones. The function returns zero on success or a negative error code if any pass signals failure.
Wiring Passes Together: The Registry Pattern
Rather than hard-coding passes in the pipeline core, [src/pipeline/registry.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/registry.c) acts as a bootstrapper that assembles the complete analysis workflow:
static pipeline_t *global_pl = NULL;
void registry_init(const char *cwd) {
global_pl = pipeline_new(cwd);
/* Register passes here */
pipeline_register_pass(global_pl, pass_semantic, NULL);
pipeline_register_pass(global_pl, pass_imports, NULL);
pipeline_register_pass(global_pl, pass_definitions, NULL);
}
registry_run() delegates to pipeline_run(), while registry_cleanup() invokes pipeline_free() to release resources. This centralized registration makes it trivial to reorder passes or inject new analysis stages without modifying the execution engine.
Implementing a Concrete Pass
Individual passes reside in separate translation units under src/pipeline/. For example, [src/pipeline/pass_semantic.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pass_semantic.c) implements a pass that extracts type information:
void pass_semantic(pipeline_t *pl, void *ctx) {
/* Access the artifact store via pl to read previous metadata */
/* Perform semantic analysis and write new artifacts */
(void)pl; (void)ctx; /* placeholder for production logic */
}
The function signature must match exactly: return type void and parameters (pipeline_t *, void *). Passes typically interact with the artifact system (declared in artifact.h) to persist their findings for downstream consumers.
Extending the Pipeline with Custom Passes
To add a new analysis stage:
- Create a new source file in
src/pipeline/(e.g.,pass_security.c). - Implement the required function signature.
- Include the header in [
registry.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/registry.c). - Register the pass using
pipeline_register_pass()beforeregistry_run()is called.
This design decouples analysis logic from orchestration, allowing domain-specific passes (security scanning, performance profiling, cross-repo linking) to coexist without changing the underlying [pipeline.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.c) implementation.
Summary
- The multi-pass indexing pipeline orchestrates sequential analysis stages over a repository stored in a
pipeline_tstructure managed bypipeline_new()andpipeline_free(). - Passes register via
pipeline_register_pass()and execute in order throughpipeline_run(), enabling downstream stages to consume artifacts from upstream stages. - The [
registry.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/registry.c) module wires concrete passes (e.g.,pass_semantic) into the global pipeline, providing a clean extension point for new analysis types.
Frequently Asked Questions
What makes the CBM pipeline "multi-pass" rather than a single analysis?
The multi-pass indexing pipeline executes a series of discrete functions in strict order, with each pass producing intermediate artifacts stored in the pipeline state. This contrasts with single-pass tools that extract all metadata in one traversal, allowing CBM to build increasingly rich metadata layers where later passes depend on earlier results.
How does the pipeline maintain state between passes?
State persists in the pipeline_t structure allocated by pipeline_new(), which includes the working directory and an expandable array of registered passes. Because each pass receives the same pipeline pointer, it can access shared artifact stores and context data left by preceding passes.
Can I add custom passes without modifying core CBM files?
You can add new passes by creating a source file implementing the required signature and registering it in [src/pipeline/registry.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/registry.c). While you must edit the registry to include your pass, you do not need to change the generic execution logic in [pipeline.c](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.c) or [pipeline.h](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.h).
What happens if a pass returns an error?
Currently, the pipeline_run() implementation returns zero unconditionally; individual passes should propagate errors through the context pointer or by setting global error flags. Production deployments typically extend the pass function signature to return error codes, allowing pipeline_run() to halt execution and return a negative value when a stage fails.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →