How to Perform Dynamic Call Tracing with Code-Graph-RAG: A Complete Guide
Code-Graph-RAG provides a built-in CallGraphTracer that hooks into Python's sys.monitoring API (PEP 669) to capture dynamic caller-callee relationships during runtime, exporting aggregated results to a JSON-Lines format for downstream graph construction.
Code-Graph-RAG ships with a sophisticated runtime call-graph tracer designed to capture dynamic execution patterns that static analysis alone cannot detect. By recording real-time caller-to-callee relationships while your Python tests, scripts, or services execute, the system builds accurate execution graphs that power retrieval-augmented generation workflows. This guide explains how to perform dynamic call tracing with Code-Graph-RAG using the production-grade CallGraphTracer class and the lightweight trace_calls utility.
Architecture of the Dynamic Call Tracing System
The tracing system operates through two distinct layers that separate instrumentation from data serialization.
The Aggregation Layer (CallGraphTracer)
At the core of the system resides the CallGraphTracer class, implemented in codebase_rag/trace/tracer.py. This component hooks into the sys.monitoring API to intercept every Python-to-Python call event (PY_START). It aggregates call pairs into an internal mapping, counts occurrences, attributes calls to specific workloads, and samples receiver types for bound methods up to the limit defined by TRACE_RECEIVER_SAMPLE_LIMIT in codebase_rag/constants.py.
The Trace Interchange Layer
Data persistence follows a standardized JSON-Lines format defined in codebase_rag/trace/records.py. The format consists of a TraceHeader followed by CallRecord objects, enabling language-agnostic downstream processing. The write_trace_file utility handles serialization, while read_trace_file supports ingestion by the graph builder in codebase_rag/graph_loader.py.
How the CallGraphTracer Works
The tracer executes a seven-stage pipeline to capture and export dynamic call graphs:
-
Initialize. The
__init__method accepts an absoluterepo_rootpath and constructs a root prefix for fast path filtering. It initializes the internal_pairsdictionary and workload registry. -
Set Workload. The
set_workloadmethod maps a human-readable name (e.g.,"unit-tests") to a numeric ID, enabling lightweight attribution of calls to specific execution contexts. -
Start Monitoring. The
startmethod registers the_on_py_startcallback withsys.monitoringforPY_STARTevents, activating the profiler. -
Filter Scope. On each event,
_in_scopevalidates that the source file resides underrepo_rootand is not within excluded directories (e.g.,venv,.git), returningFalsefor out-of-scope frames. -
Record Calls. The
_on_py_startcallback extracts caller and calleeCodeTypeobjects, increments occurrence counters, stores the current workload ID, and captures concrete receiver types for bound methods during the firstTRACE_RECEIVER_SAMPLE_LIMITencounters. -
Stop Monitoring. The
stopmethod disables monitoring events and releases the profiler ID throughsys.monitoring, ensuring clean interpreter state. -
Export. The
writemethod converts aggregated pairs intoCallRecordobjects via therecords()generator, constructs aTraceHeader, and serializes everything to JSON-Lines usingwrite_trace_filefromrecords.py.
Practical Implementation Examples
Code-Graph-RAG provides two distinct APIs for dynamic tracing depending on your performance and granularity requirements.
Full-Featured Tracing with CallGraphTracer
Use the CallGraphTracer class for production workloads requiring aggregation, workload attribution, and receiver type sampling:
from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer
def my_workload():
import mypackage.sample
mypackage.sample.do_something()
# Initialize tracer with repository root
tracer = CallGraphTracer(repo_root=Path(__file__).parent.parent)
# Optional: Name the workload for attribution
tracer.set_workload("sample-run")
# Execute tracing
tracer.start()
try:
my_workload()
finally:
tracer.stop()
# Export to JSON-Lines
trace_path = Path("dynamic_trace.jsonl")
record_count = tracer.write(trace_path)
print(f"Wrote {record_count} aggregated call records to {trace_path}")
Key implementation details:
- Only calls originating from files under
repo_rootare recorded - Receiver types are sampled for the first
TRACE_RECEIVER_SAMPLE_LIMITinvocations of each bound method - The exported file can be consumed by
read_trace_fileinrecords.pyand fed tograph_loader.py
Lightweight Tracing with trace_calls
For quick investigations or the L3 evaluation suite, use the trace_calls function from evals/calls_trace.py:
from pathlib import Path
from evals.calls_trace import trace_calls
def workload():
import mypackage.sample
mypackage.sample.do_something()
edges = trace_calls(
workload=workload,
target=Path(__file__).parent.parent,
project_name="myproject",
)
for caller, callee in sorted(edges):
print(f"{caller} → {callee}")
Unlike the full tracer, this implementation uses the classic sys.settrace API and returns a simple set of (caller, callee) qualified name tuples without aggregation or receiver type sampling.
Choosing the Right Approach
- Use
CallGraphTracerwhen tracing large production workloads, when you need workload attribution, or when building graphs for the RAG pipeline viagraph_loader.py - Use
trace_callsfor development debugging, unit tests, or when you need minimal overhead and simple edge enumeration
Core Source Files and Components
Understanding the module structure helps integrate dynamic tracing into custom workflows:
-
codebase_rag/trace/tracer.py: Contains theCallGraphTracerclass implementing thesys.monitoringhook, aggregation logic, and thewriteexport method. -
codebase_rag/trace/records.py: DefinesTraceHeader,CallRecord, and thewrite_trace_file/read_trace_fileutilities for JSON-Lines serialization. -
evals/calls_trace.py: Implements the lightweighttrace_callsfunction usingsys.settracefor evaluation scenarios and rapid prototyping. -
codebase_rag/constants.py: Houses configuration constants includingTRACE_TOOL_NAME,TRACE_RECEIVER_SAMPLE_LIMIT, and directory exclusion lists. -
codebase_rag/graph_loader.py: Consumes trace files produced by the tracer to inject dynamic edges into the knowledge graph model.
Summary
Dynamic call tracing with Code-Graph-RAG captures runtime execution patterns through Python's sys.monitoring API:
- The
CallGraphTracerclass incodebase_rag/trace/tracer.pyprovides production-grade instrumentation with call aggregation and workload attribution - Traces export to a standardized JSON-Lines format defined in
codebase_rag/trace/records.py, enabling seamless integration with the graph loader - The
trace_callsutility inevals/calls_trace.pyoffers a lightweight alternative usingsys.settracefor simpler use cases - Both approaches filter calls by repository root and exclude standard directories like
venvand.git - Receiver type sampling and workload naming support sophisticated downstream analysis in the RAG pipeline
Frequently Asked Questions
What is the difference between the CallGraphTracer and trace_calls functions?
The CallGraphTracer leverages the modern sys.monitoring API (PEP 669) for low-overhead aggregation, supports workload attribution, and samples receiver types for bound methods. In contrast, trace_calls uses the legacy sys.settrace mechanism and simply returns a set of qualified name pairs without metadata or aggregation capabilities.
How does the tracer handle calls outside the target repository?
The _in_scope method in tracer.py filters frames by verifying that the source file path starts with the configured repo_root and does not match excluded patterns defined in codebase_rag/constants.py. Calls from system libraries, virtual environments, or VCS directories are automatically discarded.
What format does the tracer use for exported data?
The tracer exports to JSON-Lines format where the first line contains a TraceHeader object and subsequent lines contain CallRecord objects. This format is defined in codebase_rag/trace/records.py and is consumed by read_trace_file and the graph_loader.py ingestion pipeline.
Can I trace multiple workloads in a single session?
Yes. Invoke set_workload() with different names between execution phases. The tracer maps each workload name to a numeric ID and attributes subsequent call observations to the active workload, enabling comparative analysis of different execution scenarios within one trace file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →