How to Use Runtime Call Tracing to Expose Dynamic Dispatch in Code-Graph-RAG

Runtime call tracing in code‑graph‑rag captures dynamic dispatch by instrumenting the Python interpreter with sys.monitoring (PEP 669), sampling concrete receiver types at call sites to reveal which actual classes are invoked at runtime.

The code‑graph‑ag repository provides a production‑ready runtime tracer that goes beyond static analysis to expose polymorphic call targets, virtual method resolution, and callback patterns. By sampling the concrete class of self or cls during live execution, you can differentiate between a static call to Base.method() and a dynamic dispatch to Derived.method().

Understanding the CallGraphTracer Architecture

The CallGraphTracer class in [codebase_rag/trace/tracer.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/tracer.py) implements the core instrumentation. It registers a callback for the PY_START monitoring event, which fires on every function entry.

Event Handling and Sampling Strategy

The tracer records three critical data points for each caller‑callee pair:

  • Call count – total invocations observed
  • Workload attribution – which named workload(s) triggered the call
  • Receiver types – the concrete class of the receiver, sampled for the first TRACE_RECEIVER_SAMPLE_LIMIT occurrences (default typically 5–10 samples)

This sampling limit prevents memory explosion while still capturing the polymorphic variants that matter for understanding dynamic dispatch patterns.

Output Format and Dynamic Dispatch Metadata

When tracer.write() persists results (lines 66‑77), each CallRecord includes a receiver_types tuple. During graph ingestion, these populate the edge properties defined in [codebase_rag/constants/trace.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants/trace.py):

  • TRACE_PROP_RECEIVER_TYPES – the dynamic_receiver_types field attached to CALL edges
  • dynamic flag – marks edges that exhibit runtime‑resolved dispatch

Step‑by‑Step: Tracing Dynamic Dispatch in Practice

1. Initialize the Tracer

from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer

repo_root = Path(__file__).resolve().parents[2]
tracer = CallGraphTracer(repo_root)

The repo_root parameter enables path normalization—call locations are stored relative to this root for portable trace files.

2. Start Monitoring

tracer.start()  # lines 74-87: registers sys.monitoring PY_START callback

This activates the low‑overhead instrumentation. The tracer now intercepts every Python function entry with minimal performance penalty compared to older sys.settrace approaches.

3. Define and Label Your Workload

def workload():
    from mypackage import MyClass
    obj = MyClass()
    obj.virtual_method()      # polymorphic dispatch site

    obj.another_method()      # may also resolve dynamically

tracer.set_workload("example_run")  # lines 63-73: tags subsequent calls

Workload labeling allows you to aggregate calls by scenario—unit tests, integration tests, or production traffic patterns.

4. Execute and Capture

workload()
tracer.stop()  # lines 89-99: unregisters callbacks, finalizes buffers

Stopping is critical—unregistered monitoring leaks can cause interpreter slowdowns.

5. Persist and Inspect

trace_file = Path("cgr-trace.jsonl")
record_count = tracer.write(trace_file)  # returns number of CallRecords written

The resulting JSON‑L contains structured call records:

{
  "kind": "call",
  "caller": {
    "path": ".../mypackage/base.py",
    "qualname": "Base.virtual_method",
    "line": 10
  },
  "callee": {
    "path": ".../mypackage/derived.py",
    "qualname": "Derived.virtual_method",
    "line": 5
  },
  "count": 3,
  "workloads": ["example_run"],
  "receiver_types": ["mypackage.derived.Derived"]
}

The receiver_types array reveals that Base.virtual_method was invoked on Derived instances—dynamic dispatch exposed.

6. Ingest Into the Graph

[codebase_rag/trace/ingest.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/ingest.py) reads the JSON‑L file and constructs graph edges:

  • Creates CALL relationships between caller and callee nodes
  • Attaches dynamic: true when receiver types differ from the static callee type
  • Populates dynamic_receiver_types with observed concrete classes

Key Configuration and Limits

Constant Location Purpose
TRACE_RECEIVER_SAMPLE_LIMIT tracer.py Max receiver type samples per caller‑callee pair
TRACE_PROP_RECEIVER_TYPES constants/trace.py Edge property key for dynamic types
TRACE_PROP_DYNAMIC constants/trace.py Boolean flag for runtime‑resolved dispatch

Adjust TRACE_RECEIVER_SAMPLE_LIMIT upward if your codebase exhibits high polymorphism (deep inheritance hierarchies, many plugin implementations).

Complementary Tracing Implementations

The repository includes additional tracers for cross‑language analysis:

These share the same ingestion pipeline, enabling unified dynamic call graphs across polyglot systems.

Performance Characteristics

  • Overhead: Typically 5–15% for CPU‑bound workloads, less for I/O‑bound code
  • Memory: Bounded by TRACE_RECEIVER_SAMPLE_LIMIT × unique caller‑callee pairs
  • Trace file size: ~200–500 bytes per CallRecord in JSON‑L format

Summary

  • Runtime call tracing via sys.monitoring (PEP 669) enables low‑overhead capture of live call graphs in code‑graph‑rag
  • CallGraphTracer in tracer.py samples concrete receiver types to expose dynamic dispatch that static analysis misses
  • receiver_types on trace records and dynamic_receiver_types on graph edges reveal which actual classes are invoked at polymorphic sites
  • Workload labeling supports scenario‑based call aggregation for targeted analysis
  • Ingestion pipeline in ingest.py converts traces into queryable graph structure with full dynamic dispatch metadata

Frequently Asked Questions

How does runtime call tracing differ from static analysis for finding virtual method calls?

Static analysis infers possible call targets from inheritance hierarchies and type annotations, often over‑approximating with many false positives. Runtime call tracing records the exact receiver class at each invocation, giving you precise dynamic dispatch information grounded in actual execution—critical for understanding plugin architectures, dependency injection, and framework callbacks where static types are deliberately abstract.

What Python versions support the tracer's monitoring approach?

The CallGraphTracer requires Python 3.12+ for native sys.monitoring support per PEP 669. For earlier Python versions, use the fallback implementation in [evals/calls_trace.py](https://github.com/vitali87/code-graph-rag/blob/main/evals/calls_trace.py) which uses sys.settrace with higher overhead and no receiver type sampling.

Can I trace multiple workloads and merge their results?

Yes. Call the set_workload() method before each distinct execution phase—the tracer accumulates all observations in memory. When you invoke write(), the workloads array on each record lists all workload labels that triggered that caller‑callee pair. The ingestion pipeline naturally merges traces; duplicate edges have their count aggregated and workloads unioned.

Why limit the number of receiver type samples?

The TRACE_RECEIVER_SAMPLE_LIMIT protects against pathological cases like tight loops over heterogenous collections that might produce thousands of distinct receiver types. In practice, the first 5–10 samples typically capture all polymorphic variants in well‑designed code; this bounded sampling keeps memory usage predictable while preserving the dynamic dispatch signal you need for graph analysis.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →