How to Integrate Production eBPF Profiles (Parca, Pyroscope) with Code‑Graph‑RAG Tracing

TLDR: Code‑Graph‑RAG ingests pprof profiles from eBPF continuous profilers like Parca or Pyroscope via the cgr trace convert command, converting production call stacks into dynamic CALLS edges that merge with the static analysis graph.

The vitali87/code-graph-rag repository provides a dedicated conversion pipeline that bridges the gap between runtime eBPF profiling and static code visualization. By converting gzipped pprof binaries into a JSON‑L trace format, you can augment your repository's call graph with real execution paths observed in production.

What are eBPF Production Profiles?

eBPF continuous profilers such as Parca and Pyroscope use kernel‑level instrumentation to sample CPU stacks with minimal overhead. These tools export data in the standard pprof binary format (protocol buffer‑encoded, gzip‑compressed), which contains:

  • Sample timestamps and values
  • Stack frame addresses and symbol names
  • Mapping information for shared libraries

Code‑Graph‑RAG treats these profiles as dynamic traces—evidence of actual function invocations that may differ from static call‑graph predictions due to polymorphism, callbacks, or runtime plugin loading.

The Three‑Step Integration Workflow

1. Export a pprof Profile from Your Profiler

Both Parca and Pyroscope expose HTTP endpoints for profile retrieval.

Parca endpoint pattern:

curl -L "https://parca.example.com/api/v1/profile?query=my-service&from=now-5m&to=now&format=pprof" \
     -o prod_profile.pb.gz

Pyroscope endpoint pattern:

curl -L "http://pyroscope-server:4040/render?query=my-service.cpu&from=now-5m&to=now&format=pprof" \
     -o prod_profile.pb.gz

The resulting file is a gzipped protobuf message conforming to the pprof profile schema.

2. Convert the Profile with cgr trace convert

Code‑Graph‑RAG's CLI provides the convert subcommand under the trace namespace. According to the implementation in [codebase_rag/trace/cli.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/cli.py), this entry point delegates to convert_ebpf_pprof in [codebase_rag/trace/ebpf_pprof.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/ebpf_pprof.py).

CLI conversion:

cgr trace convert prod_profile.pb.gz \
    --format ebpf \
    --repo-path /path/to/your/repo \
    --output dynamic_trace.jsonl

Parameter reference:

Flag Description
profile (positional) Path to the gzipped pprof file
--format ebpf Required to trigger the eBPF‑specific parser
--repo-path Root of the repository for path normalization
--output Destination for the JSON‑L trace file

3. Merge the Dynamic Trace with Your Static Graph

The converted trace contains line‑delimited JSON objects representing observed call stacks. Load these into your existing graph with:

cgr trace load \
    --graph graph.json \
    --trace dynamic_trace.jsonl \
    --output merged_graph.json

How the Conversion Works Internally

The convert_ebpf_pprof function in codebase_rag/trace/ebpf_pprof.py performs these operations:

  1. Parse the protobuf using the vendored schema_pb2.py definitions
  2. Iterate samples and reconstruct call stacks from sample.location_id chains
  3. Normalize paths against the provided --repo-path to enable source linking
  4. Emit JSON‑L records with schema: {"caller": "funcA", "callee": "funcB", "source": "ebpf", "weight": N}

Each record includes a weight derived from sample values (CPU time or sample count), allowing downstream analysis to prioritize hot paths.

Python API for Programmatic Integration

For custom pipelines, use the Python API directly:

from codebase_rag.trace.ebpf_pprof import convert_ebpf_pprof
from codebase_rag.graph_loader import load_graph, apply_trace
from pathlib import Path

# Step 1: Convert eBPF profile to trace format

profile_path = Path("prod_profile.pb.gz")
repo_root = Path("/path/to/your/repo")
trace_path = Path("dynamic_trace.jsonl")

sample_count = convert_ebpf_pprof(
    profile_path,
    repo_path=repo_root,
    output_path=trace_path,
    format="ebpf",
)
print(f"Converted {sample_count} samples to {trace_path}")

# Step 2: Load and merge with static graph

graph = load_graph("graph.json")
graph = apply_trace(graph, trace_path, trace_source="ebpf")
graph.save("merged_graph.json")

The apply_trace function distinguishes dynamic edges by tagging them with source: "ebpf", enabling filtered queries like: "Show me all static calls that were never observed in production."

Handling Symbolization and Path Mapping

Production eBPF profiles may contain raw addresses rather than symbol names if the profiler lacks access to debug symbols. The conversion pipeline assumes server‑side symbolization (Parca/Pyroscope's default) but provides hooks for:

  • Offline symbol resolution via addr2line or llvm-symbolizer
  • Path remapping for reproducible builds where source paths differ between compilation and analysis environments

Configure remapping via environment variable or CLI flag as documented in [docs/guide/dynamic-tracing.md](https://github.com/vitali87/code-graph-rag/blob/main/docs/guide/dynamic-tracing.md):

CGR_PATH_MAP="/build/path:/repo/path" cgr trace convert ...

CI/CD Integration Example

Automate profile ingestion in your deployment pipeline:


# .github/workflows/profile-ingest.yml

name: Ingest Production Profiles

on:
  schedule:
    - cron: '0 */6 * * *'  # Every 6 hours

jobs:
  ingest:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      
      - name: Install code-graph-rag
        run: pip install code-graph-rag
      
      - name: Fetch latest Parca profile
        run: |
          curl -L "${{ secrets.PARCA_URL }}/api/v1/profile?query=prod-api&from=now-6h&format=pprof" \
               -o profile.pb.gz
      
      - name: Convert and merge
        run: |
          cgr trace convert profile.pb.gz --format ebpf --repo-path . --output trace.jsonl
          cgr trace load --graph base_graph.json --trace trace.jsonl --output graph.json
      
      - name: Upload enriched graph
        uses: actions/upload-artifact@v4
        with:
          name: enriched-graph
          path: graph.json

Summary

  • Export pprof profiles from Parca, Pyroscope, or OpenTelemetry eBPF profilers via their HTTP APIs
  • Convert using cgr trace convert --format ebpf, implemented in codebase_rag/trace/ebpf_pprof.py
  • Merge the resulting JSON‑L trace with your static graph using cgr trace load or the Python API
  • Distinguish dynamic eBPF edges from static analysis via the source metadata field
  • Automate ingestion in CI pipelines for continuously updated production‑informed call graphs

Frequently Asked Questions

What eBPF profilers are officially supported by Code‑Graph‑RAG?

The ebpf_pprof.py converter handles any profiler emitting standard pprof format, including Parca, Pyroscope, Grafana Cloud Profiles, and the OpenTelemetry eBPF profiler. The key requirement is gzip‑compressed protobuf output conforming to the pprof schema.

How do I handle profiles with unsymbolized addresses?

Ensure your profiler performs server‑side symbolization—both Parca and Pyroscope support uploading debug info or fetching symbols from symbol servers. For offline processing, pre‑process profiles with pprof CLI tools to resolve addresses before conversion.

Can I weight edges by CPU time versus sample count?

Yes. The converter reads the sample.value array and uses the default sample type (typically CPU nanoseconds or samples) as the edge weight. The protobuf sample_type field determines which value is extracted; this matches pprof's semantic conventions.

What's the performance impact of converting large production profiles?

Conversion is streaming and memory‑bounded: ebpf_pprof.py processes samples iteratively without loading the entire profile into RAM. Typical throughput exceeds 100,000 samples/second on modern hardware. For multi‑gigabyte profiles, consider the --samples-limit flag to cap ingestion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →