# How to Enable Dynamic Tracing for Python with Code-Graph-RAG

> Learn how to enable dynamic tracing for Python using Code-Graph-RAG's CallGraphTracer. Capture caller-callee relationships, dynamic dispatch, and workload attribution with this powerful tool.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-04

---

**Code-Graph-RAG provides a `CallGraphTracer` that hooks into Python's `sys.monitoring` API (PEP 669) to record every Python-to-Python call within a repository root, capturing caller-callee relationships, dynamic dispatch types, and workload attribution in a JSON trace format.**

Code-Graph-RAG (CG-RAG) ships with a high-performance runtime instrumentation system for Python that leverages the low-overhead `sys.monitoring` hooks introduced in PEP 669. The `CallGraphTracer` class in [`codebase_rag/trace/tracer.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/tracer.py) enables you to capture detailed call graphs including virtual method dispatch patterns and test-specific execution paths. This article explains the two primary methods to enable dynamic tracing for Python with Code-Graph-RAG: via the built-in pytest plugin or directly through the Python API.

## Architecture of the Tracing System

The tracing architecture centers on three core components that work together to capture and serialize call graph data.

### Core Implementation Files

- **`CallGraphTracer`** ([`codebase_rag/trace/tracer.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/tracer.py)): Registers a `PY_START` callback with `sys.monitoring` using the tool ID defined in [`codebase_rag/constants/trace.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/trace.py) as `TRACE_TOOL_NAME = "cgr-trace"`. It filters events to the repository root, aggregates caller-to-callee pairs, and samples receiver types for dynamic dispatch analysis.

- **Pytest Plugin** ([`codebase_rag/trace/pytest_plugin.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/pytest_plugin.py)): Provides the `--cgr-trace` command-line integration that automatically manages tracer lifecycle, workload attribution, and trace file generation at session completion.

- **Trace Constants** ([`codebase_rag/constants/trace.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/trace.py)): Defines configuration values including the tracer tool name, default output locations, and formatting constants used throughout the tracing subsystem.

### Key Capabilities

The tracer captures **workload attribution** via `set_workload()`, allowing you to tag every observed call with a specific test identifier or job name. It also performs **receiver-type sampling** when the callee's first argument matches a known receiver name (e.g., `self`), recording the concrete type for the first few occurrences to analyze dynamic dispatch patterns.

## Method 1: Enable Tracing via the Pytest Plugin

The most common workflow uses the built-in pytest plugin, which activates automatically when you add the `--cgr-trace` flag. The plugin lazily imports the tracer only for traced runs, avoiding overhead for normal test executions.

```bash
pytest -q --cgr-trace --cgr-trace-output=trace.json

```

When you invoke pytest with `--cgr-trace`, the plugin performs the following operations internally:

1. Creates a `CallGraphTracer` instance scoped to the repository root
2. Calls `tracer.start()` to register the `PY_START` event with `sys.monitoring.use_tool_id(..., cs.TRACE_TOOL_NAME)`
3. Sets the current test's node ID as the workload via `tracer.set_workload(item.nodeid)` before each test
4. Clears the workload via `tracer.set_workload(None)` after each test completes
5. Calls `tracer.stop()` and `tracer.write(output_path)` during `pytest_sessionfinish` to flush the aggregated data

The resulting [`trace.json`](https://github.com/vitali87/code-graph-rag/blob/main/trace.json) contains structured entries including caller and callee file paths, qualified names, execution counts, associated workloads, and sampled receiver types:

```json
{
  "header": { "tracer": "cgr-trace" },
  "records": [
    {
      "caller": { "path": ".../module_a.py", "qualname": "ClassA.method_a", "line": 10 },
      "callee": { "path": ".../module_b.py", "qualname": "ClassB.method_b", "line": 5 },
      "count": 3,
      "workloads": ["test_module.py::TestFoo::test_bar"],
      "receiver_types": ["module_b.ClassB"]
    }
  ]
}

```

## Method 2: Enable Tracing Programmatically

For ad-hoc scripts, CI pipelines, or non-pytest environments, instantiate `CallGraphTracer` directly and manage the lifecycle manually.

```python
from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer

# 1. Initialize the tracer with your repository root

repo_root = Path(__file__).parent.parent
tracer = CallGraphTracer(repo_root=repo_root)

# 2. Start monitoring (registers the sys.monitoring callback)

tracer.start()

# 3. Assign a workload name for downstream filtering

tracer.set_workload("my_script_run")

# 4. Execute the code you want to trace

import my_package.some_module as sm
sm.do_something()  # Dynamic dispatch on 'self' will be captured here

# 5. Stop tracing and persist results

tracer.stop()
output_file = Path("dynamic_trace.json")
record_count = tracer.write(output_file)
print(f"Wrote {record_count} call records to {output_file}")

```

This approach is useful for tracing arbitrary subprocesses or application entry points. For example, wrapping a CI command:

```python
import subprocess
from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer

def run_with_trace(command: list[str], repo_root: Path, out: Path) -> None:
    tracer = CallGraphTracer(repo_root)
    tracer.start()
    tracer.set_workload("ci_job_" + "_".join(command))
    subprocess.run(command, check=True)
    tracer.stop()
    tracer.write(out)

run_with_trace(
    ["python", "-m", "my_app.main"],
    repo_root=Path.cwd(),
    out=Path("ci_trace.json")
)

```

## Understanding Trace Data Features

### Dynamic Dispatch Capture

When the tracer observes a call where the first argument matches a receiver pattern (such as `self`), it samples the concrete type of that argument for the first few occurrences of the call pair. This sampling logic in `_on_py_start` (lines 28-41 of [`tracer.py`](https://github.com/vitali87/code-graph-rag/blob/main/tracer.py)) captures virtual method calls and trait object patterns without recording every single invocation.

### Workload Attribution

The `set_workload()` method stores a workload ID for every observed call, enabling downstream analysis to filter call graphs by specific test cases, CI jobs, or logical execution units. This provenance tracking is essential for associating runtime behavior with specific code changes or test failures.

## Summary

- **Code-Graph-RAG** enables dynamic tracing through the `CallGraphTracer` class in [`codebase_rag/trace/tracer.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/tracer.py), which integrates with Python's `sys.monitoring` API using the tool name `cgr-trace`.

- **Two activation methods** are available: the `--cgr-trace` pytest flag for test suites, or direct instantiation of `CallGraphTracer` for scripts and CI pipelines.

- **Workload attribution** via `set_workload()` allows you to tag call records with test identifiers or job names for targeted analysis.

- **Receiver-type sampling** captures dynamic dispatch information by recording concrete types of method receivers during execution.

- **Trace output** uses a JSON interchange format containing call records with caller/callee metadata, execution counts, workloads, and receiver types, written via the `write()` method.

## Frequently Asked Questions

### What Python version is required for Code-Graph-RAG tracing?

Code-Graph-RAG requires Python 3.12 or later, as it depends on the `sys.monitoring` API introduced in PEP 669. This API provides the low-overhead, interpreter-wide hooks necessary for efficient runtime tracing without modifying bytecode.

### How does the tracer impact runtime performance?

The tracer uses `sys.monitoring` with a `PY_START` event callback, which introduces minimal overhead compared to traditional `sys.settrace` approaches. The implementation filters events by repository root early in the callback chain and aggregates data in memory, ensuring the instrumentation is suitable for large test suites and production workloads.

### Can I trace code outside of a pytest environment?

Yes. While the `--cgr-trace` pytest plugin provides convenient automation, you can instantiate `CallGraphTracer` directly in any Python script, import the modules you wish to trace, and call `start()`, `stop()`, and `write()` manually to capture execution traces without a testing framework.

### What information does the trace output file contain?

The trace file contains a `TraceHeader` followed by an array of `CallRecord` objects. Each record includes the caller and callee file paths, qualified names, line numbers, execution counts, associated workloads (set via `set_workload()`), and sampled receiver types for dynamic dispatch analysis.