How to Integrate Dead Code Analysis into CI Pipelines: A Complete Guide
Integrate dead code analysis into CI pipelines by installing the cgr CLI from the vitali87/code-graph-rag repository and executing cgr dead-code --fail-on-found to automatically block builds containing unreachable symbols.
The vitali87/code-graph-rag repository provides a language-agnostic dead code detection engine that identifies unreachable code through call graph traversal. Located in codebase_rag/dead_code.py, the engine analyzes CALLS and REFERENCES relationships to distinguish live code from dead symbols. Adding this analysis to your continuous integration workflow ensures every commit ships only reachable production code, reducing maintenance burden and technical debt accumulation.
Understanding the Dead Code Detection Architecture
Graph Extraction and Root Identification
The dead code engine operates on a call graph stored in Memgraph or an in-memory representation. The collect_dead_code function in codebase_rag/dead_code.py extracts nodes and relationships, then identifies root symbols—entry points automatically considered reachable regardless of incoming calls.
Root detection leverages language-specific helpers including _is_root, _is_rust_runtime_root, _is_js_well_known_symbol_root, and _is_c_cpp_entry_root. These functions recognize framework patterns such as Rust #[test] attributes, Python decorators, React lifecycle methods, and C++ main functions. This multi-language support enables the engine to work with Python, Rust, JavaScript/TypeScript, C/C++, Java, and C# codebases within a single pipeline.
Reachability Analysis via BFS
After identifying roots, the engine executes _walk to perform a multi-source BFS across CALLS and REFERENCES edges. The dead_code_from_graph function orchestrates this traversal, tracking visited symbols while respecting language semantics including method overrides and factory patterns. Any symbol not visited during the walk is flagged as dead and formatted for output via _emit_dead_code and _build_dead_code_table.
Installing the Tool in Your Pipeline
Install the package directly from your repository checkout. The pyproject.toml defines the cgr entry point, enabling immediate CLI access:
pip install .
For development environments or faster iteration, use an editable install:
uv pip install -e .
Configuring Dead Code Analysis for CI
The Typer-based CLI in codebase_rag/cli.py exposes the dead-code command through the _dead_code_config function. Configure the analysis to match your project's structure by specifying entry points and exclusion patterns:
cgr dead-code \
--project-name myapp \
--entry-point myapp.main:run \
--decorator-root @app.route \
--exclude "*generated*" \
--format json \
--output dead_code_report.json
Key configuration options include:
--project-name: The namespace prefix used in the graph database--entry-point: Additional entry points beyond auto-detected roots--decorator-root: Custom decorators signifying reachable code (e.g.,@task,@controller)--exclude: Glob patterns to skip generated files or directories--include-tests/--no-include-tests: Control test code analysis (defaults toTrue)--classes: Enable class-level roots for languages where constructors serve as entry points
Failing Builds on Dead Code Detection
To enforce code hygiene, add the --fail-on-found flag. This triggers a non-zero exit code when unreachable symbols are detected, immediately blocking the pipeline:
cgr dead-code --project-name myapp --fail-on-found
Complete GitHub Actions Integration Example
The following workflow installs the tool, runs analysis, and uploads the report as a build artifact:
name: CI
on:
push:
branches: [main]
pull_request:
jobs:
dead-code:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install cgr
run: pip install .
- name: Run dead-code analysis
run: |
cgr dead-code \
--project-name myapp \
--entry-point myapp.main:run \
--exclude "*generated*" \
--format json \
--output dead_code_report.json \
--fail-on-found
- name: Upload report
if: always()
uses: actions/upload-artifact@v4
with:
name: dead-code-report
path: dead_code_report.json
This configuration ensures builds fail immediately upon detecting dead code while preserving the JSON report for forensic analysis or downstream notification steps.
Programmatic Integration with Python
For advanced customization, bypass the CLI and invoke the engine directly using collect_dead_code and default_dead_code_config:
from codebase_rag.dead_code import collect_dead_code, default_dead_code_config
from codebase_rag.graph_loader import load_graph
graph = load_graph(project_name="myapp")
config = default_dead_code_config(
include_tests=False,
include_classes=True,
exclude_patterns=("*/generated/*", "*/migrations/*"),
)
dead_rows = collect_dead_code(graph.ingestor, "myapp", config)
for row in dead_rows:
print(f"Dead: {row['qualified_name']} ({row['path']}:{row['start_line']})")
This approach allows dynamic configuration of the DeadCodeConfig object, enabling runtime exclusion rules or custom root detection logic without modifying the CLI source in codebase_rag/cli.py.
Managing False Positives and Generated Code
Dead code analysis requires careful handling of generated files to prevent false positives. Use exclusion patterns to filter protocol buffers, GraphQL generated clients, or ORM migrations:
cgr dead-code \
--project-name myapp \
--exclude "*generated*" \
--exclude "*/migrations/*" \
--exclude "*.pb.go"
The engine's deterministic output depends solely on the graph structure, making CI caching straightforward when the source snapshot remains unchanged.
Summary
- Install via pip: The
cgrCLI installs directly from the repository usingpip install .as defined inpyproject.toml - Use
--fail-on-found: Non-zero exit codes transform dead code detection into a CI gatekeeper - Configure exclusions: Apply
--excludepatterns to handle generated code and prevent false positives from files likeclient/core/* - Language agnostic: The engine supports multiple languages through specialized root detectors in
codebase_rag/dead_code.py - Flexible reporting: Generate JSON for programmatic consumption or human-readable tables for build logs using
_build_dead_code_table
Frequently Asked Questions
How does the dead code engine determine which code is reachable?
The engine performs a multi-source BFS starting from root symbols detected by functions like _is_root and _is_rust_runtime_root in codebase_rag/dead_code.py. It traverses CALLS and REFERENCES edges in the graph, marking all visited symbols as live. Any symbol not reached during this traversal—executed by the _walk function—is reported as dead.
Can I run dead code analysis without a live Memgraph database?
Yes. The collect_dead_code function accepts any graph ingestor implementing the expected interface, including in-memory test harnesses. For CI pipelines, ensure your workflow either starts a Memgraph service container or loads a serialized graph snapshot using the appropriate loader from codebase_rag/graph_loader.
How do I prevent the CI pipeline from failing on specific dead code I want to keep?
While the engine does not currently support inline ignore annotations, you can use the --exclude flag with glob patterns to skip specific files or directories. For fine-grained control, use the Python API to wrap collect_dead_code and implement custom filtering logic before checking for dead symbols.
Does the tool support monorepos with multiple projects?
Yes. The --project-name parameter isolates analysis to specific namespaces within the graph. Run separate cgr dead-code commands for each project in your monorepo, each with the appropriate project name and entry points defined in _dead_code_config, to generate distinct reports for each codebase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →