How Dead Code Detection Works with Degree Filtering in codebase-memory-mcp

Dead code detection in codebase-memory-mcp operates by constructing a symbol-usage graph and iteratively pruning nodes with zero in-degree until reaching a fixed point, effectively identifying unreachable definitions without expensive data-flow analysis.

The codebase-memory-mcp repository implements a fast, language-agnostic dead code detection system using a degree filtering approach on a symbol-usage graph. By analyzing the connectivity between definitions and their references, the tool efficiently identifies unreachable code across diverse codebases. This article examines the C-engine implementation in DeusData/codebase-memory-mcp, graph construction mechanics, and practical API usage for detecting dead symbols.

Building the Symbol-Usage Graph

The detection process begins with graph construction inside the C-engine located in internal/cbm/cbm.c. While scanning the source tree, the engine creates a node for every definition—functions, classes, variables—and adds directed edges representing every use of that symbol, including calls, imports, and references.

Two specialized extraction modules handle the discovery phase:

This construction phase produces a directed graph where an edge from node A to node B indicates that symbol A depends on or references symbol B.

Degree Filtering Algorithm

Once the graph is built, the system calculates connectivity metrics for each node. The cbm_node_t structure defined in internal/cbm/cbm.h tracks two critical counters:

  • In-degree – The count of other nodes that reference this definition.
  • Out-degree – The count of symbols this node references.

The dead-code detector applies degree filtering by iteratively removing nodes whose in-degree is 0 and that are not designated entry points (such as main, exported public symbols, or symbols marked with "keep" annotations).

Iterative Pruning to Fixed Point

The algorithm runs until convergence, producing a cascade effect that identifies transitive dead code. The core logic in internal/cbm/cbm.c implements the following pseudocode:

while (true) {
    bool changed = false;
    for each node n in graph {
        if (n.in_degree == 0 && !is_root(n)) {
            remove_node(n);                // drop the node and its edges
            decrement_out_degrees_of_adjacent_nodes(n);
            changed = true;
        }
    }
    if (!changed) break;                  // fix-point reached
}

When a node is removed, the system decrements the in-degree of every node it pointed to. This reduction may cause additional nodes to drop to zero in-degree, triggering their removal in subsequent iterations. The process repeats until no new dead symbols are identified, ensuring that only symbols reachable (directly or indirectly) from root entry points remain in the final graph.

Implementation in the C Core

The degree filtering approach is intentionally lightweight, operating in linear time relative to the number of edges. Because the underlying graph is language-agnostic, the detector works across programming languages without requiring language-specific parsers for the analysis phase.

Eliminated nodes are reported as dead code with their file locations and removal reasons. The detector outputs a JSON list containing dead symbols, their source files, and line numbers, enabling integration with CI pipelines and code editors.

Detecting Dead Code via Python API and CLI

The Python wrapper in pkg/pypi/src/codebase_memory_mcp/__init__.py exposes the detect_dead_code() method, which invokes the C-core degree filtering routine:

from codebase_memory_mcp import CodebaseMemoryMCP

# Initialise the MCP with a path to the repository

mcp = CodebaseMemoryMCP(root_path="/path/to/repo")

# Run the analysis (this builds the graph and applies degree filtering)

dead_symbols = mcp.detect_dead_code()

# `dead_symbols` is a list of dictionaries:

# [{'name': 'unused_helper', 'file': 'src/util.c', 'line': 42}, ...]

for sym in dead_symbols:
    print(f"Dead: {sym['name']} → {sym['file']}:{sym['line']}")

For command-line usage, the CLI implemented in pkg/pypi/src/codebase_memory_mcp/_cli.py provides direct access:


# Scan the repository and output dead symbols as JSON

codebase-memory-mcp --root /path/to/repo --detect-dead-code --output dead.json

Both interfaces leverage the same degree-filtering engine, ensuring consistent results across programmatic and command-line workflows.

Summary

  • Graph Construction – The C-engine in internal/cbm/cbm.c builds a symbol-usage graph using extract_defs.c for definitions and extract_calls.c for references.
  • Degree Metrics – Each node tracks in_degree and out_degree via the cbm_node_t structure in internal/cbm/cbm.h.
  • Filtering Logic – Nodes with zero in-degree that are not entry points are iteratively removed until a fixed point is reached, creating a cascade effect that catches transitive dead code.
  • Performance – The algorithm runs in linear time relative to edge count, making it scalable for large repositories.
  • Integration – Python and CLI interfaces in pkg/pypi/src/codebase_memory_mcp/ provide convenient access to the core detection engine.

Frequently Asked Questions

What is degree filtering in dead code detection?

Degree filtering is a graph-based technique that identifies dead code by analyzing the connectivity (in-degree) of symbols in a usage graph. Symbols with zero in-degree have no incoming references, meaning no other part of the codebase depends on them, making them candidates for removal unless they are entry points like main or public exports.

How does codebase-memory-mcp distinguish entry points from dead code?

The algorithm checks each candidate node using an is_root() function that identifies entry points such as main functions, exported public symbols, or symbols explicitly marked with "keep" annotations. These root nodes are preserved regardless of their in-degree count to ensure the application remains executable.

Why does the detection algorithm use iterative fixed-point iteration instead of a single pass?

A single pass would miss transitive dead code—symbols that become unreachable only after their dependents are removed. By iterating until no changes occur (fixed point), the algorithm captures cascade effects where removing one dead symbol reduces the in-degree of others, causing them to become dead in subsequent iterations.

Which source files contain the core degree filtering implementation?

The core logic resides in internal/cbm/cbm.c, which orchestrates the graph building and filtering loop. The node structure definitions are in internal/cbm/cbm.h, while symbol discovery happens in internal/cbm/extract_defs.c and internal/cbm/extract_calls.c. Python bindings are located in pkg/pypi/src/codebase_memory_mcp/__init__.py and pkg/pypi/src/codebase_memory_mcp/_cli.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →