How Hybrid LSP Semantic Type Resolution Works in Codebase Memory MCP

Hybrid LSP semantic type resolution is a lightweight C implementation that runs atop tree-sitter ASTs to compute IDE-grade type information and call edges without requiring external language-server processes.

The DeusData/codebase-memory-mcp repository builds a persistent knowledge graph by parsing source files with tree-sitter, then applying a second Hybrid LSP pass that implements type-resolution algorithms similar to tsserver, pyright, and rust-analyzer. This pass enriches the graph with fully-qualified call targets and confidence scores, enabling fast structural queries directly from SQLite storage.

The Six-Phase Hybrid LSP Pipeline

After tree-sitter generates the initial AST, the Hybrid LSP engine executes six sequential phases for each of the nine supported language families. These phases are implemented per-language in internal/cbm/lsp/.

Phase 1: Import Binding

The pipeline begins by collecting all import statements and populating a scoped name-to-type map. In internal/cbm/lsp/py_lsp.c, the function py_lsp_bind_imports handles both import X and from X import Y syntax, inserting imported symbols into the current CBMScope before statement processing begins.

Phase 2: Statement Processing

The walker traverses top-level statements—assignments, control flow blocks, class definitions, and function declarations—registering variables, fields, and methods via py_process_statement. This establishes the initial symbol table for the current file scope.

Phase 3: Expression Type Evaluation

For every expression node, the system computes a concrete type using py_eval_expr_type. This covers literals, binary operations, and function calls, producing a CBMType that feeds into downstream attribute lookups.

Phase 4: Attribute Lookup

Attribute chains such as obj.attr.method are resolved using language-specific rules. Python implements MRO (Method Resolution Order) traversal via py_lookup_attribute, while TypeScript handles generic substitution and Rust manages trait-method dispatch, all falling back to the CBMTypeRegistry for built-in types.

Phase 5: Call Resolution

The core py_resolve_calls_in function recursively walks the AST, emitting a CBMResolvedCall for every call site where the callee can be resolved to a fully-qualified name. Each entry includes a confidence score—1.0 for direct bindings, less than 1.0 for heuristic matches—that persists as a RESOLVED_CALLS edge in the knowledge graph.

Phase 6: Cross-File Linking

In the final phase, the indexer uses collected import maps to resolve re-exports, wildcard imports, and intra-package references (designated Phase 9 in the source). This links call edges across compilation units, completing the graph without requiring a live LSP server connection.

Core Data Structures

All phases share three fundamental structures defined in internal/cbm/lsp/type_registry.c:

  • CBMScope: A hierarchical symbol table providing cbm_scope_* functions for name-to-type bindings.
  • CBMTypeRegistry: Pre-computed tables of standard-library types stored in internal/cbm/lsp/generated/*_stdlib_data.c files.
  • CBMResolvedCallArray: A growable list of resolved call edges that becomes persistent graph data.

Language-Specific Implementation

Each supported language follows the same binding-to-resolution pattern, differing only in semantic rules:

Auto-generated files under internal/cbm/lsp/generated/ (e.g., python_stdlib_data.c) provide the CBMTypeRegistry data for each language's standard library.

Working with the Hybrid LSP API

The following C example demonstrates initializing the Hybrid LSP context for Python and extracting resolved calls:

/* Allocate an arena for temporary allocations */
CBMArena *arena = cbm_arena_new();

/* Load source text (normally read from a file) */
const char *src = "import numpy as np; np.array([1,2])";
int src_len = strlen(src);

/* Registry contains stdlib types for Python (generated data) */
extern const CBMTypeRegistry python_registry;
CBMResolvedCallArray calls = {0};

/* Initialise the LSP context */
PyLSPContext ctx;
py_lsp_init(&ctx, arena, src, src_len,
            &python_registry,
            "__main__", &calls);

/* Register imports extracted by the tree-sitter pass */
py_lsp_add_import(&ctx, "np", "numpy");
py_lsp_bind_imports(&ctx);

/* Walk the AST and emit resolved calls */
TSNode root = ts_tree_root_node(ctx.tree);
py_resolve_calls_in(&ctx, root);

/* `calls` now contains fully-qualified callee names */
for (size_t i = 0; i < calls.len; ++i) {
    printf("call %zu: %s (conf %.2f)\n",
           i, calls.items[i].callee_qn, calls.items[i].confidence);
}

After indexing, query the persistent SQLite graph for resolved call edges:

codebase-memory-mcp query_graph '
MATCH (f:Function)-[:RESOLVED_CALLS]->(g:Function)
WHERE f.file = "src/main.py"
RETURN f.name, g.name, g.module
' | jq .

Summary

  • Hybrid LSP runs as a C-based layer atop tree-sitter, eliminating the need for external language-server processes during indexing.
  • The six-phase pipeline handles import binding, statement processing, expression evaluation, attribute lookup, call resolution, and cross-file linking.
  • Core data structures (CBMScope, CBMTypeRegistry, CBMResolvedCallArray) provide language-agnostic symbol management.
  • Language-specific implementations in internal/cbm/lsp/py_lsp.c, ts_lsp.c, and similar files adapt the pipeline to Python, TypeScript, Go, Rust, and other supported languages.
  • Results persist as RESOLVED_CALLS edges in the SQLite-backed MCP graph, enabling fast structural queries for AI coding agents.

Frequently Asked Questions

What is the difference between Hybrid LSP and a traditional language server?

Hybrid LSP is a static analysis pass that computes type resolution once during indexing and stores results in a persistent graph, whereas traditional language servers like pyright or tsserver run continuously in the background to provide IDE features. The Hybrid approach enables fast queries without maintaining a live server process.

How does Hybrid LSP handle cross-file imports?

During the Cross-File Linking phase, the system uses collected import maps from all indexed files to resolve re-exports, wildcard imports, and intra-package references. This creates RESOLVED_CALLS edges that point to fully-qualified names across the entire codebase, not just within a single file.

Which programming languages support Hybrid LSP resolution?

The implementation supports nine language families: Python, TypeScript, JavaScript, JSX, TSX, PHP, C#, Go, C/C++, Java, Kotlin, and Rust. Each has a dedicated implementation file in internal/cbm/lsp/ (e.g., go_lsp.c, rust_lsp.c) that adapts the six-phase pipeline to language-specific semantics.

Where is the type resolution data stored?

Resolution data persists in the CBMTypeRegistry structures and the final SQLite-backed knowledge graph. Standard library types are pre-computed in generated files like internal/cbm/lsp/generated/python_stdlib_data.c, while project-specific call edges are stored as RESOLVED_CALLS relationships in the graph database.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →