How to Perform Dynamic Call Tracing with Code-Graph-RAG: A Complete Guide

Code-Graph-RAG provides a built-in CallGraphTracer that hooks into Python's sys.monitoring API (PEP 669) to capture dynamic caller-callee relationships during runtime, exporting aggregated results to a JSON-Lines format for downstream graph construction.

Code-Graph-RAG ships with a sophisticated runtime call-graph tracer designed to capture dynamic execution patterns that static analysis alone cannot detect. By recording real-time caller-to-callee relationships while your Python tests, scripts, or services execute, the system builds accurate execution graphs that power retrieval-augmented generation workflows. This guide explains how to perform dynamic call tracing with Code-Graph-RAG using the production-grade CallGraphTracer class and the lightweight trace_calls utility.

Architecture of the Dynamic Call Tracing System

The tracing system operates through two distinct layers that separate instrumentation from data serialization.

The Aggregation Layer (CallGraphTracer)

At the core of the system resides the CallGraphTracer class, implemented in codebase_rag/trace/tracer.py. This component hooks into the sys.monitoring API to intercept every Python-to-Python call event (PY_START). It aggregates call pairs into an internal mapping, counts occurrences, attributes calls to specific workloads, and samples receiver types for bound methods up to the limit defined by TRACE_RECEIVER_SAMPLE_LIMIT in codebase_rag/constants.py.

The Trace Interchange Layer

Data persistence follows a standardized JSON-Lines format defined in codebase_rag/trace/records.py. The format consists of a TraceHeader followed by CallRecord objects, enabling language-agnostic downstream processing. The write_trace_file utility handles serialization, while read_trace_file supports ingestion by the graph builder in codebase_rag/graph_loader.py.

How the CallGraphTracer Works

The tracer executes a seven-stage pipeline to capture and export dynamic call graphs:

  1. Initialize. The __init__ method accepts an absolute repo_root path and constructs a root prefix for fast path filtering. It initializes the internal _pairs dictionary and workload registry.

  2. Set Workload. The set_workload method maps a human-readable name (e.g., "unit-tests") to a numeric ID, enabling lightweight attribution of calls to specific execution contexts.

  3. Start Monitoring. The start method registers the _on_py_start callback with sys.monitoring for PY_START events, activating the profiler.

  4. Filter Scope. On each event, _in_scope validates that the source file resides under repo_root and is not within excluded directories (e.g., venv, .git), returning False for out-of-scope frames.

  5. Record Calls. The _on_py_start callback extracts caller and callee CodeType objects, increments occurrence counters, stores the current workload ID, and captures concrete receiver types for bound methods during the first TRACE_RECEIVER_SAMPLE_LIMIT encounters.

  6. Stop Monitoring. The stop method disables monitoring events and releases the profiler ID through sys.monitoring, ensuring clean interpreter state.

  7. Export. The write method converts aggregated pairs into CallRecord objects via the records() generator, constructs a TraceHeader, and serializes everything to JSON-Lines using write_trace_file from records.py.

Practical Implementation Examples

Code-Graph-RAG provides two distinct APIs for dynamic tracing depending on your performance and granularity requirements.

Use the CallGraphTracer class for production workloads requiring aggregation, workload attribution, and receiver type sampling:

from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer

def my_workload():
    import mypackage.sample
    mypackage.sample.do_something()

# Initialize tracer with repository root

tracer = CallGraphTracer(repo_root=Path(__file__).parent.parent)

# Optional: Name the workload for attribution

tracer.set_workload("sample-run")

# Execute tracing

tracer.start()
try:
    my_workload()
finally:
    tracer.stop()

# Export to JSON-Lines

trace_path = Path("dynamic_trace.jsonl")
record_count = tracer.write(trace_path)
print(f"Wrote {record_count} aggregated call records to {trace_path}")

Key implementation details:

  • Only calls originating from files under repo_root are recorded
  • Receiver types are sampled for the first TRACE_RECEIVER_SAMPLE_LIMIT invocations of each bound method
  • The exported file can be consumed by read_trace_file in records.py and fed to graph_loader.py

Lightweight Tracing with trace_calls

For quick investigations or the L3 evaluation suite, use the trace_calls function from evals/calls_trace.py:

from pathlib import Path
from evals.calls_trace import trace_calls

def workload():
    import mypackage.sample
    mypackage.sample.do_something()

edges = trace_calls(
    workload=workload,
    target=Path(__file__).parent.parent,
    project_name="myproject",
)

for caller, callee in sorted(edges):
    print(f"{caller} → {callee}")

Unlike the full tracer, this implementation uses the classic sys.settrace API and returns a simple set of (caller, callee) qualified name tuples without aggregation or receiver type sampling.

Choosing the Right Approach

  • Use CallGraphTracer when tracing large production workloads, when you need workload attribution, or when building graphs for the RAG pipeline via graph_loader.py
  • Use trace_calls for development debugging, unit tests, or when you need minimal overhead and simple edge enumeration

Core Source Files and Components

Understanding the module structure helps integrate dynamic tracing into custom workflows:

  • codebase_rag/trace/tracer.py: Contains the CallGraphTracer class implementing the sys.monitoring hook, aggregation logic, and the write export method.

  • codebase_rag/trace/records.py: Defines TraceHeader, CallRecord, and the write_trace_file/read_trace_file utilities for JSON-Lines serialization.

  • evals/calls_trace.py: Implements the lightweight trace_calls function using sys.settrace for evaluation scenarios and rapid prototyping.

  • codebase_rag/constants.py: Houses configuration constants including TRACE_TOOL_NAME, TRACE_RECEIVER_SAMPLE_LIMIT, and directory exclusion lists.

  • codebase_rag/graph_loader.py: Consumes trace files produced by the tracer to inject dynamic edges into the knowledge graph model.

Summary

Dynamic call tracing with Code-Graph-RAG captures runtime execution patterns through Python's sys.monitoring API:

  • The CallGraphTracer class in codebase_rag/trace/tracer.py provides production-grade instrumentation with call aggregation and workload attribution
  • Traces export to a standardized JSON-Lines format defined in codebase_rag/trace/records.py, enabling seamless integration with the graph loader
  • The trace_calls utility in evals/calls_trace.py offers a lightweight alternative using sys.settrace for simpler use cases
  • Both approaches filter calls by repository root and exclude standard directories like venv and .git
  • Receiver type sampling and workload naming support sophisticated downstream analysis in the RAG pipeline

Frequently Asked Questions

What is the difference between the CallGraphTracer and trace_calls functions?

The CallGraphTracer leverages the modern sys.monitoring API (PEP 669) for low-overhead aggregation, supports workload attribution, and samples receiver types for bound methods. In contrast, trace_calls uses the legacy sys.settrace mechanism and simply returns a set of qualified name pairs without metadata or aggregation capabilities.

How does the tracer handle calls outside the target repository?

The _in_scope method in tracer.py filters frames by verifying that the source file path starts with the configured repo_root and does not match excluded patterns defined in codebase_rag/constants.py. Calls from system libraries, virtual environments, or VCS directories are automatically discarded.

What format does the tracer use for exported data?

The tracer exports to JSON-Lines format where the first line contains a TraceHeader object and subsequent lines contain CallRecord objects. This format is defined in codebase_rag/trace/records.py and is consumed by read_trace_file and the graph_loader.py ingestion pipeline.

Can I trace multiple workloads in a single session?

Yes. Invoke set_workload() with different names between execution phases. The tracer maps each workload name to a numeric ID and attributes subsequent call observations to the active workload, enabling comparative analysis of different execution scenarios within one trace file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →