Static Analysis vs Dynamic Tracing in Code-Graph-RAG: Implementation Guide

Static analysis parses source files to extract guaranteed call relationships without execution, while dynamic tracing instruments running programs to record actual call paths including dynamic dispatch and reflection.

Code-Graph-RAG builds comprehensive call graphs by combining static analysis and dynamic tracing to power retrieval-augmented generation over codebases. While static analysis provides fast, precise baseline coverage by walking ASTs, dynamic tracing fills the gaps by observing runtime behavior across Python, JVM, Node.js, .NET, and other languages. Understanding when to use each approach ensures you capture both guaranteed dependencies and runtime-specific execution paths.

How Static Analysis Works

Static analysis in Code-Graph-RAG operates by parsing source files into abstract syntax trees (ASTs) and extracting calls that can be resolved without type inference or execution.

The implementation lives in codebase_rag/evals/static_calls.py. The oracle walks the AST, builds a parent map, and resolves imports via the _import_map structure. It creates edges only for ast.Call nodes whose func is an ast.Name (see lines 45-52 and 80-86). This means it captures direct calls to imported functions (e.g., from module import foo) and top-level functions within the same module, but excludes method calls, attribute accesses, and dynamic dispatch.

Edges produced through static analysis receive the attribute dynamic: false, indicating they represent guaranteed static relationships.

How Dynamic Tracing Works

Dynamic tracing executes the program—typically the test suite or a specific workload—and records actual call relationships observed at runtime.

The Python implementation resides in codebase_rag/trace/python_tracer.py and utilizes sys.monitoring for low-overhead instrumentation. You can invoke it via pytest --cgr-trace or programmatically using the CallGraphTracer class. For other languages, the system supports JVM (byte-code instrumenting agent), Node.js (V8 sampling profiler), .NET (dotnet-trace), C/C++ (-finstrument-functions via cgr_trace_shim.c), PHP, Lua, Dart, Go, and eBPF.

When ingested via cgr trace ingest, dynamic edges receive dynamic: true and additional properties including dynamic_call_count, dynamic_workloads, and dynamic_receiver_types.

Coverage Characteristics and Trade-offs

Static analysis guarantees high precision with no false positives, but offers low recall. It only captures calls that are statically certain, omitting method invocations, reflective calls, monkey-patching, and registry-based patterns.

Dynamic tracing provides high recall for any code path actually exercised during the traced workload. However, it is inherently limited by coverage—calls never executed remain unseen, and you must ensure your test suite or workload exercises the relevant code paths.

Runtime Performance Impact

Static analysis imposes zero runtime cost on your application; it operates as a quick AST walk over source files.

Dynamic tracing overhead varies significantly by language and instrumentation method:

  • Python: Uses sys.monitoring with minimal overhead suitable for production.
  • JVM: Byte-code instrumenting agent adds approximately microseconds per call.
  • Node.js: V8 sampling profiler produces minimal impact through sampling.
  • .NET: dotnet-trace sampling approach keeps overhead low.
  • C/C++: -finstrument-functions provides exact tracing but with heavier performance impact.

Practical Implementation Examples

Running Static Analysis

To evaluate static call recall against your repository:

python -m codebase_rag.evals.static_calls \
    --target /path/to/your/repo \
    --project-name myproject \
    --out-dir ./static-results

This uses the oracle in static_calls.py (lines 97-135) to compute statically-certain call edges.

Recording Python Traces via pytest

cd /path/to/your/repo
pytest --cgr-trace               # writes cgr-trace.jsonl

cgr trace ingest cgr-trace.jsonl --repo-path .

The pytest plugin remains inert unless the --cgr-trace flag is present.

Programmatic Python Tracing

from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer

tracer = CallGraphTracer(Path("/path/to/your/repo"))
tracer.start()
try:
    run_your_workload()          # any Python code you want to profile

finally:
    tracer.stop()
tracer.write(Path("my-trace.jsonl"))

import subprocess
subprocess.run(["cgr", "trace", "ingest", "my-trace.jsonl",
                "--repo-path", "/path/to/your/repo"])

Tracing JVM Workloads

make jvm-agent                      # creates build/cgr-jvm-agent.jar

java -javaagent:build/cgr-jvm-agent.jar="include=com.example;repo=/path/to/repo;output=cgr-trace.jsonl" \
     -jar my-tests.jar

cgr trace ingest cgr-trace.jsonl --repo-path /path/to/repo

Summary

  • Static analysis in codebase_rag/evals/static_calls.py provides fast, precise baseline graphs by parsing ASTs for guaranteed calls, marking edges with dynamic: false.
  • Dynamic tracing instruments running code via language-specific agents (Python sys.monitoring, JVM byte-code, C -finstrument-functions) to capture runtime-only relationships with dynamic: true and additional metadata.
  • Combine both approaches to achieve high precision on static guarantees while filling coverage gaps through runtime observation.
  • Use static analysis for quick sanity checks and large-scale codebase scans; use dynamic tracing to validate assumptions, discover reflective patterns, and profile production workloads.

Frequently Asked Questions

Can I use static analysis and dynamic tracing together in the same project?

Yes. Code-Graph-RAG is designed to merge both data sources. Static analysis provides a fast baseline graph, while dynamic tracing ingests additional edges to capture runtime-specific behavior. The cgr trace ingest command merges traces into your existing graph, with edges tagged by their source type.

Why does static analysis miss some function calls?

The static oracle in static_calls.py intentionally conservative: it only creates edges when ast.Call nodes contain ast.Name as the function identifier (lines 45-52). This excludes method calls on objects (obj.method()), attribute accesses, and calls resolved through dynamic dispatch or import tricks. These limitations ensure zero false positives but require dynamic tracing to capture the full picture.

How do I trace performance-sensitive production code?

For production environments, use sampling-based tracers or low-overhead instrumentation. Python's sys.monitoring implementation in codebase_rag/trace/python_tracer.py adds minimal overhead. For systems languages, consider eBPF or the sampling modes available for Node.js and .NET rather than the exact instrumentation used for C/C++ testing.

Where are dynamic traces ingested in the graph database?

After generating a trace file (e.g., cgr-trace.jsonl), the cgr trace ingest command (implemented in codebase_rag/trace/trace_cmd.py) merges the observed calls into your repository's graph. Dynamic edges receive additional attributes including call counts and receiver types, enabling downstream queries to weight frequently-used paths higher in RAG contexts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →