How to Enable Dynamic Call Tracing in Python with Code-Graph-RAG

Use Code-Graph-RAG's CallGraphTracer class to capture Python-to-Python function calls at runtime via the sys.monitoring API, then export the results to a JSON-L trace file for graph analysis.

Dynamic call tracing bridges the gap between static code analysis and actual runtime behavior. The Code-Graph-RAG project provides a production-ready tracer that records real function invocations, enabling you to build accurate call graphs that reflect how your code actually executes. This guide walks through activation methods, implementation details, and practical workflows using the official source code.

Architecture Overview: Three-Stage Runtime Collection

The tracer operates through distinct phases implemented in codebase_rag/trace/tracer.py. Understanding these stages helps troubleshoot issues and optimize collection performance.

Stage 1: Setup and Registration

The CallGraphTracer.start() method registers a profiling tool with Python's sys.monitoring infrastructure (PEP 669). This installs a callback for the PY_START event, which fires before every Python function executes.

Only one tracer can be active per interpreter—sys.monitoring enforces a single profiler tool registration, making the tracer non-reentrant.

Stage 2: In-Memory Collection

On each PY_START event, the _on_py_start callback performs three critical validations:

  • Verifies both caller and callee files reside under the repository root
  • Excludes directories listed in codebase_rag/constants/trace.py (.venv, node_modules, etc.)
  • Updates a counter for the (caller, callee) pair in an internal dictionary

For bound methods, the tracer optionally samples receiver types during the first few invocations—controlled by TRACE_RECEIVER_SAMPLE_LIMIT—enabling dynamic dispatch resolution in downstream analysis.

Stage 3: Export to Trace Format

The CallGraphTracer.write() method transforms collected edges into CallRecord objects, attaches workload metadata through a TraceHeader, and serializes to .jsonl format using write_trace_file from codebase_rag/trace/records.py.

Installation and Quick Start

Ensure Code-Graph-RAG is installed in your environment:

pip install codebase-rag

The tracer requires Python 3.12+ for full sys.monitoring support, though partial compatibility exists for earlier versions using sys.setprofile as a fallback.

Method 1: Zero-Configuration Tracing with Pytest

The pytest plugin in codebase_rag/trace/pytest_plugin.py provides the fastest path to dynamic call collection.

cd /path/to/your-repo
pytest --cgr-trace

This command automatically:

  • Instantiates CallGraphTracer with the current working directory as repo_root
  • Starts tracing before test collection
  • Stops and writes cgr-trace.jsonl after teardown

Customize the output location:

pytest --cgr-trace --cgr-trace-output=./traces/integration.jsonl

Method 2: Programmatic Tracing for Custom Workloads

For scripts, services, or interactive sessions, instantiate CallGraphTracer directly from codebase_rag/trace/tracer.py.

from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer

# Initialize with absolute path to repository root

tracer = CallGraphTracer(Path("/path/to/your-repo"))

# Label this run for downstream filtering

tracer.set_workload("data-pipeline-ingestion")

# Execute traced code with guaranteed cleanup

tracer.start()
try:
    process_customer_batch()      # Your application code here

    validate_output()
finally:
    tracer.stop()

# Export and verify

trace_path = Path("cgr-trace.jsonl")
record_count = tracer.write(trace_path)
print(f"Exported {record_count} call records to {trace_path}")

The set_workload() method converts string labels to integer IDs via internal hashing, storing the label in TraceHeader for workload-specific graph reconstruction.

Method 3: Context Manager Wrapper

For cleaner integration, wrap CallGraphTracer in a context manager that handles lifecycle automatically.

from pathlib import Path
from codebase_rag.trace.tracer import CallGraphTracer

class traced:
    """Context manager for scoped dynamic call tracing."""
    
    def __init__(self, repo_root: Path, workload: str | None = None):
        self.tracer = CallGraphTracer(repo_root)
        if workload:
            self.tracer.set_workload(workload)
        self._output_path = Path("cgr-trace.jsonl")

    def __enter__(self):
        self.tracer.start()
        return self.tracer

    def __exit__(self, exc_type, exc, tb):
        self.tracer.stop()
        self.tracer.write(self._output_path)
        return False  # Don't suppress exceptions

# Usage: trace a specific operation block

with traced(Path("/my/repo"), workload="batch-job-v3") as t:
    initialize_workers()
    run_computation()
    aggregate_results()

Scope Filtering and Configuration

The _in_scope() method in codebase_rag/trace/tracer.py applies three exclusion layers:

  1. Directory patterns — .venv, node_modules, __pycache__, .git (defined in codebase_rag/constants/trace.py)
  2. Repository boundary — files outside repo_root are silently skipped
  3. File extensions — only .py files are processed (with C extension calls recorded as <native>)

Modify exclusions by editing EXCLUDED_DIRS in codebase_rag/constants/trace.py or subclassing CallGraphTracer:

from codebase_rag.trace.tracer import CallGraphTracer

class CustomTracer(CallGraphTracer):
    def _in_scope(self, filepath: Path) -> bool:
        # Add custom logic before delegating to parent

        if "generated" in filepath.parts:
            return False
        return super()._in_scope(filepath)

Receiver Sampling for Polymorphism Analysis

Dynamic dispatch complicates static analysis. The tracer addresses this by capturing concrete receiver types during early invocations of bound methods.

Configuration in codebase_rag/constants/trace.py:

TRACE_RECEIVER_SAMPLE_LIMIT = 3  # Capture type for first 3 calls per method

The receiver_types field in CallRecord contains a frequency map of observed types, enabling graph algorithms to weight edges by actual runtime behavior rather than declared types.

Trace Output Format

The .jsonl file produced by write_trace_file follows this structure (defined in codebase_rag/trace/records.py):

  • Header line: TraceHeader with tool_name, version, timestamp, and workload_id
  • Record lines: CallRecord objects with caller, callee, count, and optional receiver_types

Example processing:

import json
from pathlib import Path

trace_path = Path("cgr-trace.jsonl")
with trace_path.open() as f:
    header = json.loads(f.readline())
    for line in f:
        record = json.loads(line)
        print(f"{record['caller']} → {record['callee']}: {record['count']}")

Performance Considerations

  • Overhead: Typically 5-15% for CPU-bound workloads; higher for call-heavy code
  • Memory: Unbounded growth during tracing—long-running processes should chunk output via periodic write() calls
  • Startup cost: First call to start() compiles regex patterns for exclusion matching; consider warming in production deployments

Summary

  • Use pytest --cgr-trace for immediate, configuration-free tracing of test suites
  • Instantiate CallGraphTracer programmatically for scripts, services, and custom workloads
  • Apply set_workload() to tag runs for downstream filtering and comparison
  • Verify scope filtering through _in_scope() behavior—only repository-internal Python calls are captured
  • Export with tracer.write() to generate standard JSON-L for Code-Graph-RAG ingestion

Frequently Asked Questions

What Python versions support Code-Graph-RAG dynamic call tracing?

Python 3.12+ provides full functionality via sys.monitoring (PEP 669). Earlier versions fall back to sys.setprofile with reduced precision—C calls and some built-in invocations may be missed. The tracer detects capabilities at import time in codebase_rag/trace/tracer.py.

Can I trace multiple repositories simultaneously?

No—sys.monitoring restricts active profilers to one per interpreter. However, you can instantiate multiple CallGraphTracer objects sequentially, each with different repo_root values, and merge their output files. For microservice architectures, consider distributed tracing integration instead.

How do I exclude specific modules or functions from tracing?

Directory-level exclusion is configured in codebase_rag/constants/trace.py via EXCLUDED_DIRS. For finer granularity, subclass CallGraphTracer and override _in_scope() or _on_py_start() to implement custom filtering logic. Function-level exclusion is not natively supported due to sys.monitoring API limitations.

Where is the trace file format documented?

The canonical specification resides in codebase_rag/trace/records.py, where CallRecord, FramePoint, and TraceHeader dataclasses define the schema. Additional user documentation is available in docs/guide/dynamic-tracing.md within the repository, covering ingestion workflows and integration with the RAG pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →