How codebase-memory-mcp's Hybrid LSP Achieves Type Resolution Without External Language Servers
The codebase-memory-mcp (CBM) project embeds language servers directly into its binary and merges their output with static analysis through an in-process cache, eliminating the need for developers to install or manage external LSP processes.
The Hybrid LSP architecture in codebase-memory-mcp combines deterministic static analysis with Language Server Protocol (LSP) precision by vending popular language servers (TypeScript, Python, etc.) directly into the repository. This design allows CBM to resolve types by communicating with these embedded servers over STDIO while maintaining a unified type graph that merges baseline Tree-sitter analysis with LSP-derived type information.
The Three-Stage Hybrid Resolution Pipeline
CBM's type resolution operates through a tightly integrated three-stage pipeline that processes source files and produces a merged type graph.
Stage 1: Collect and Cache LSP Diagnostics
When the pipeline encounters a source file, CBM initializes a virtual LSP session for that specific language. The system sends standard LSP requests—including textDocument/hover, textDocument/typeDefinition, and textDocument/definition—to the embedded language server, then stores the returned type information in an in-process LSP cache.
The cache structures and public API are defined in src/pipeline/lsp_resolve.h. This header declares the functions for querying cached LSP data and configuring resolution parameters, ensuring that subsequent look-ups for the same symbol are O(1) and incur no additional IPC overhead.
Stage 2: Traverse the AST with Hybrid Resolution
During the AST walk, CBM's baseline resolver (powered by Tree-sitter parsers) first attempts static type resolution. When the static view is ambiguous or missing, the system queries the LSP cache for the same symbol via the hybrid resolver implemented in internal/cbm/lsp_all.c.
This file orchestrates the decision logic that determines when to fall back to LSP data. Each cached LSP response includes a confidence score derived from the LSP's symbol kind and response certainty. The resolver uses this score to decide whether the LSP type should override the static analysis result.
Stage 3: Merge and Prune Type Graphs
The final stage merges the baseline type graph with the LSP-provided graph, using a confidence floor to prevent low-quality LSP data from degrading analysis. Only LSP types exceeding the CBM_LSP_CONFIDENCE_FLOOR (default 0.6) replace baseline types.
The merging process utilizes iterator utilities defined in internal/cbm/lsp/lsp_node_iter.h to walk LSP-provided symbol trees efficiently. The merged graph then feeds into downstream pipeline stages such as call-graph construction and memory analysis.
Why External Language Servers Aren't Required
CBM eliminates external dependencies through three architectural decisions:
-
Embedded LSP Binaries: The repository vendors compiled language servers in directories like
vendored/tsserverandvendored/pyright. CBM launches these as child processes within the same address space, communicating over STDIO rather than requiring socket-based network connections. -
In-Process Caching: All LSP responses remain in memory for the duration of the analysis run. The cache eliminates redundant IPC calls and provides constant-time lookups for repeated symbol queries.
-
Unified Resolution API: CBM exposes a single function
cbm_lsp_resolve(symbol_id, out_type)declared insrc/pipeline/lsp_resolve.h. Call sites throughout the codebase request type information without needing to know whether the result originated from static analysis or the LSP cache.
Additionally, the entire LSP layer can be disabled at runtime by setting the environment variable CBM_LSP_DISABLED, allowing users to fall back to pure static analysis for benchmarking or debugging purposes.
Implementation Examples
Initializing the Hybrid LSP Pipeline
#include "src/pipeline/lsp_resolve.h"
int main(int argc, char **argv) {
/* Spin up embedded LSP child processes */
cbm_lsp_initialize();
/* Run the analysis pipeline */
cbm_pipeline_run(argv[1]);
/* Clean up language server processes */
cbm_lsp_shutdown();
return 0;
}
The initialization and shutdown functions are declared in src/pipeline/lsp_resolve.h.
Querying Types Through the Hybrid Resolver
#include "src/pipeline/lsp_resolve.h"
CBMType *resolve_symbol_type(const char *symbol_name) {
CBMSymbolId sid = cbm_lookup_symbol(symbol_name);
CBMType *type = NULL;
/* Attempt static resolution first */
if (!cbm_static_resolve(sid, &type)) {
/* Fallback to LSP cache */
cbm_lsp_resolve(sid, &type);
}
return type;
}
The hybrid resolution logic is implemented in internal/cbm/lsp_all.c.
Adjusting the Confidence Threshold
#include "src/pipeline/lsp_resolve.h"
int main(void) {
/* Only accept LSP types with >= 0.8 confidence */
cbm_set_lsp_confidence_floor(0.8);
cbm_pipeline_run("project_path");
return 0;
}
Confidence floor configuration is available in src/pipeline/lsp_resolve.h.
Disabling LSP for Static-Only Analysis
#include "src/pipeline/lsp_resolve.h"
int main(void) {
/* Disable LSP via environment variable */
cbm_setenv("CBM_LSP_DISABLED", "1", 1);
cbm_pipeline_run("benchmark_project");
return 0;
}
This pattern appears in the test suite files tests/test_ts_lsp.c and tests/test_py_lsp.c.
Summary
- Embedded Architecture: codebase-memory-mcp vendors language servers directly in
vendored/directories, launching them as child processes to avoid external installation requirements. - Three-Stage Pipeline: The system collects LSP diagnostics into an in-process cache, walks the AST with a hybrid resolver that falls back to LSP data when static analysis is insufficient, and merges type graphs using confidence scoring.
- Confidence-Based Merging: Only LSP types exceeding the
CBM_LSP_CONFIDENCE_FLOOR(default 0.6) override baseline analysis, ensuring quality control. - Zero-Configuration API: The unified
cbm_lsp_resolve()function provides a single interface for type resolution, abstracting whether the source was static or LSP-derived. - Runtime Control: The
CBM_LSP_DISABLEDenvironment variable andcbm_set_lsp_confidence_floor()function allow users to toggle LSP integration and adjust quality thresholds without recompiling.
Frequently Asked Questions
How does codebase-memory-mcp handle multiple languages without external servers?
CBM embeds language-specific LSP binaries for each supported language (TypeScript via vendored/tsserver, Python via vendored/pyright, etc.) within the repository. When analyzing a file, the system spawns the appropriate embedded server as a child process and communicates over STDIO. This design allows multi-language projects to receive precise type resolution without requiring developers to install Node.js, Python packages, or other runtime dependencies for language servers.
What happens when the LSP returns uncertain type information?
The system assigns a confidence score to each LSP response based on the symbol kind and response certainty. During the merge stage in internal/cbm/lsp_all.c, CBM compares this score against the CBM_LSP_CONFIDENCE_FLOOR (default 0.6). If the LSP's confidence falls below this threshold, the system retains the baseline static analysis type rather than accepting the ambiguous LSP data, preventing low-quality types from polluting the analysis graph.
Can I disable the Hybrid LSP if I only want static analysis?
Yes. Set the environment variable CBM_LSP_DISABLED to "1" before running the pipeline, or call cbm_setenv("CBM_LSP_DISABLED", "1", 1) programmatically. This is demonstrated in the test files tests/test_ts_lsp.c and tests/test_py_lsp.c. When disabled, CBM skips the LSP initialization and relies solely on its Tree-sitter-based static resolver, which is useful for benchmarking or when working in environments where spawning child processes is restricted.
Where is the LSP type cache stored during analysis?
The LSP cache is maintained entirely in memory within the process heap. The cache structures are defined in src/pipeline/lsp_resolve.h and persist for the duration of the analysis run. This in-memory approach eliminates disk I/O bottlenecks and allows O(1) lookups for repeated symbol queries, though the cache is destroyed when cbm_lsp_shutdown() is called or the process terminates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →