How DeusData Uses Codebase Memory MCP for Real-Time AI Code Intelligence
DeusData implements codebase memory MCP as a high-performance client/server architecture that maintains a live SQLite graph of code symbols, enabling AI agents to query repository structure in milliseconds via JSON-RPC over STDIO.
The DeusData codebase memory MCP provides AI coding agents with instant, semantic awareness of repository structure without requiring full re-parsing. According to the DeusData/codebase-memory-mcp source code, this system combines a native C engine, platform-specific file watchers, and a Python wrapper to deliver incremental code intelligence.
Architecture Overview
The implementation consists of three tightly-coupled layers that work together to provide real-time code analysis.
Native Engine Layer
The core functionality resides in src/main.c, a high-performance C binary that parses source code and stores symbol relationships in a compact SQLite database. This engine handles the heavy lifting of building and querying the code graph, supporting languages including C, C++, Python, and TypeScript.
File Watcher Daemon
Located in src/watcher/watcher.c, the file watcher uses platform-specific APIs—inotify on Linux and ReadDirectoryChangesW on Windows—to monitor the workspace for changes. When files are added, renamed, or edited, the watcher triggers incremental re-indexing rather than full rescans.
Python Wrapper and CLI
The pkg/pypi/src/codebase_memory_mcp/_cli.py module provides the user-facing interface. This wrapper automatically downloads pre-built binaries from GitHub releases (as seen in __init__.py lines 6-11) and exposes both a command-line interface and a Python API for embedding in applications.
Launching the MCP Daemon
DeusData designed the system for drop-in deployment. After installing via pip install codebase-memory-mcp, users initialize a workspace daemon:
codebase-memory-mcp daemon --workspace /path/to/repo
This command performs three critical operations:
- Downloads the correct native binary for the current platform
- Starts the SQLite-backed engine and spawns the file watcher
- Opens a persistent STDIO channel for JSON-RPC communication
The daemon creates a .cbm.db file in the workspace root, which stores the graph representation of the codebase.
Indexing and Graph Storage
When the daemon initializes, the engine walks the --workspace directory and constructs a semantic graph database containing three primary node types:
- File nodes: Store
path,mtime, andsizemetadata - Symbol nodes: Capture
name,kind(function, class, variable), and sourcelocation - Edge relationships: Track
defines,references, andimportsbetween symbols
This graph structure enables sophisticated queries such as locating all call sites of a specific function or identifying symbols exported from a particular module. Because the data persists in SQLite, subsequent daemon restarts can perform incremental loading rather than full rebuilding.
Incremental Updates via File Watcher
The src/watcher/watcher.h interface defines how the system maintains index freshness without re-parsing the entire repository. Upon detecting filesystem changes:
- The affected file undergoes targeted re-parsing
- Stale symbol and edge rows are purged from the database
- New relationships are inserted
- An incremental snapshot transmits to the engine via the RPC channel
This architecture ensures that long-running AI agents operate against up-to-date codebase representations, with updates reflecting in the graph within milliseconds of file saves.
Querying from Python Applications
Clients communicate with the daemon through a thin JSON-RPC layer that writes requests to the daemon's STDIO pipe. The pkg/pypi/src/codebase_memory_mcp/__init__.py exposes a main entry point that handles daemon lifecycle management:
from codebase_memory_mcp import main
# Initialize daemon for workspace
daemon = main(['daemon', '--workspace', '/my/project'])
# Query for symbol definitions
result = daemon.query({
"action": "find_definitions",
"symbol": "UserService"
})
Supported query actions include:
find_definitions: Locates where a symbol is declaredfind_references: Returns all usage sites of a symbolfind_completions: Provides context-aware autocomplete suggestions
Because the engine runs as compiled C code, round-trip latency remains in the millisecond range even for enterprise-scale repositories.
Integration with AI Coding Agents
DeusData's AI agents embed the Python wrapper to enable sophisticated code intelligence features. When a user opens a workspace, the agent launches the daemon and maintains it as a background service for the session duration.
Common integration patterns include:
- Autocomplete: Calling
find_completionsat cursor positions to suggest relevant symbols - Go-to-definition: Executing
find_definitionsto resolve token origins instantly - Refactoring: Using
find_referencesto locate all sites requiring modification before applying rename operations
This approach eliminates the need for AI agents to implement their own parsers or maintain expensive in-memory AST representations.
Practical Implementation Examples
Starting the Daemon Programmatically
For custom tool integration, spawn the daemon via subprocess and communicate via JSON-RPC:
import subprocess
import json
import time
# Launch daemon process
proc = subprocess.Popen(
['codebase-memory-mcp', 'daemon', '--workspace', '/my/app'],
stdin=subprocess.PIPE,
stdout=subprocess.PIPE,
text=True
)
# Allow indexing time
time.sleep(2)
def rpc(request: dict) -> dict:
"""Send JSON-RPC request and return parsed response."""
proc.stdin.write(json.dumps(request) + '\n')
proc.stdin.flush()
return json.loads(proc.stdout.readline())
# Find AuthToken definitions
response = rpc({
"action": "find_definitions",
"symbol": "AuthToken"
})
Command-Line Workflow
For ad-hoc analysis without Python scripting:
# Initialize workspace index
codebase-memory-mcp init --workspace /my/project
# Query references from CLI
codebase-memory-mcp query \
--action find_references \
--symbol getUserData
Embedding in LLM Applications
The wrapper integrates cleanly into agent architectures:
from codebase_memory_mcp import main as mcp_main
class CodeAssistant:
def __init__(self, root_path):
self.mcp = mcp_main(['daemon', '--workspace', root_path])
def locate_symbol(self, name):
return self.mcp.query({
"action": "find_definitions",
"symbol": name
})
Summary
- Three-layer architecture: Native C engine (
src/main.c), file watcher (src/watcher/watcher.c), and Python wrapper (pkg/pypi/src/codebase_memory_mcp/_cli.py) provide separation of concerns and optimal performance. - Real-time synchronization: Platform-specific file watchers enable incremental updates without full rescans, maintaining live accuracy for AI agents.
- Millisecond query latency: JSON-RPC over STDIO to a compiled C backend ensures fast responses even for complex graph traversals in large repositories.
- Language agnostic: Supports C, C++, Python, TypeScript, and additional languages through the unified symbol graph stored in SQLite (
.cbm.db).
Frequently Asked Questions
What storage format does codebase memory MCP use?
The system stores code graphs in a compact SQLite database named .cbm.db within the workspace root. This file contains tables for file nodes, symbol nodes, and relationship edges (defines, references, imports), enabling fast local queries without external database dependencies.
How does DeusData handle file changes without full re-indexing?
The src/watcher/watcher.c module implements platform-specific filesystem notifications (inotify on Linux, ReadDirectoryChangesW on Windows). When changes occur, only the affected file undergoes re-parsing. The engine then performs differential updates to the SQLite graph—removing stale entries and inserting new ones—while the daemon continues serving queries uninterrupted.
Which programming languages does the system support?
According to the source analysis in src/main.c, the parser handles C, C++, Python, and TypeScript out of the box. The graph structure itself is language-agnostic, storing generic symbol kinds (functions, classes, variables) and relationships that enable cross-language reference tracking in polyglot repositories.
How do AI agents establish communication with the daemon?
Agents use JSON-RPC over STDIO pipes. The Python wrapper in pkg/pypi/src/codebase_memory_mcp/__init__.py manages the daemon subprocess and provides a query() method that serializes requests to JSON, writes them to the daemon's stdin, and parses the JSON response from stdout. This simple protocol allows any language capable of spawning processes to integrate with the MCP server.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →