How to Specify Entry Points for Dead Code Analysis in Code-Graph-RAG

In the vitali87/code-graph-rag repository, entry points for dead code analysis are specified via the --entry-points CLI option in evals/dead_code.py, which accepts JSON files, YAML files, or comma-separated lists of fully-qualified symbol names.

The Code-Graph-RAG project provides a call-graph-based dead code evaluator that determines code liveness by traversing reachability from specified execution roots. Configuring dead code analysis entry points correctly ensures the tool distinguishes between truly dead code and implicitly invoked modules. The entry point resolution logic sits at the boundary between the CLI driver and the core graph engine.

Specifying Entry Points via the CLI

The evaluation script evals/dead_code.py exposes the --entry-points argument to declare which functions, classes, or scripts serve as the live roots of your application. You can supply these seeds in three interoperable formats.

JSON File Format

Provide a path to a JSON file containing an array of fully-qualified symbol strings. The parser validates the JSON structure and converts each entry into a symbol reference that the graph engine can resolve.

[
  "my_app.cli.main",
  "my_app.service.startup",
  "my_app.utils.helper"
]

YAML File Format

You may alternatively supply a YAML file with an equivalent list structure. The CLI loader detects the .yaml or .yml extension and parses the document into the internal symbol format before passing it to the analysis engine.

- my_app.cli.main
- my_app.service.startup
- my_app.utils.helper

Inline Comma-Separated List

For ad-hoc analysis or CI integration, pass the symbols directly as a comma-separated string without spaces. This bypasses file I/O and is parsed immediately by the argument handler in evals/dead_code.py.

--entry-points "my_pkg.main,another_pkg.start"

How Entry Points Drive the Analysis

Once parsed by the CLI layer, the entry point list is forwarded to codebase_rag/dead_code.py, which implements the core dead code detection logic. This module constructs a set of seed nodes from the supplied fully-qualified symbol names (e.g., package.module.Class.method).

The engine then performs a forward reachability analysis across the imported function and method call graph. Beginning from the seed nodes, it traverses the transitive closure of all possible execution paths. Any symbol visited during this traversal is marked as live code, while all unvisited symbols are reported as dead code in the final output.

Practical Usage Examples

The following examples demonstrate complete command invocations using different entry point specification methods.

Using a JSON Configuration File

Create entry_points.json enumerating your application's public API, then pass it to the evaluator:

python -m evals.dead_code \
    --repo-root /path/to/repo \
    --entry-points entry_points.json \
    --output dead_code_report.json

Using a YAML Configuration File

For human-readable configuration management, use the YAML format:

python -m evals.dead_code \
    --repo-root /path/to/repo \
    --entry-points entry_points.yaml \
    --output dead_code_report.json

Inline Symbol Specification

Run a quick scan without creating intermediate files:

python -m evals.dead_code \
    --repo-root /path/to/repo \
    --entry-points "my_app.cli.main,my_app.service.startup" \
    --output dead_code_report.json

Summary

  • Specify entry points using the --entry-points argument when invoking evals/dead_code.py from the command line
  • Supported formats include JSON arrays, YAML lists, or inline comma-separated fully-qualified symbol names
  • Symbol resolution requires fully-qualified paths (e.g., package.module.function) to correctly map to graph nodes
  • Core implementation resides in codebase_rag/dead_code.py, which builds seed nodes and executes forward reachability analysis
  • Output generation produces a list of unreachable (dead) symbols and optionally writes a structured JSON report for further processing

Frequently Asked Questions

What file formats are accepted for dead code analysis entry points?

The --entry-points option accepts JSON files containing symbol arrays, YAML files with equivalent list structures, or raw comma-separated strings passed directly on the command line. The parser in evals/dead_code.py introspects the argument to determine whether to treat it as a file path or as an inline list.

Where is the entry point processing logic implemented?

CLI argument parsing and file loading occur in evals/dead_code.py, which normalizes inputs into a standard list of symbol references. The graph traversal logic that consumes these entry points is implemented in codebase_rag/dead_code.py, where the engine constructs seed nodes and computes reachability.

How does the tool distinguish between live and dead code?

After resolving entry point strings to specific graph nodes, codebase_rag/dead_code.py performs a forward traversal from these seeds across the call graph. Any function, method, or class reachable through this traversal is classified as live; all disconnected nodes are flagged as dead code in the final report.

Can I use examples/graph_export_example.py to visualize the live code paths?

Yes, the examples/graph_export_example.py script demonstrates how to export the constructed call graph, which is useful for visualizing the reachability results from your entry point analysis alongside the dead code report.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →