# How to Integrate Production eBPF Profiles (Parca, Pyroscope) with Code‑Graph‑RAG Tracing

> Learn to integrate production eBPF profiles from Parca or Pyroscope with Code-Graph-RAG tracing. Convert call stacks into dynamic CALLS edges for enhanced analysis.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-20

---

**TLDR:** Code‑Graph‑RAG ingests pprof profiles from eBPF continuous profilers like **Parca** or **Pyroscope** via the `cgr trace convert` command, converting production call stacks into dynamic `CALLS` edges that merge with the static analysis graph.

The `vitali87/code-graph-rag` repository provides a dedicated conversion pipeline that bridges the gap between runtime eBPF profiling and static code visualization. By converting gzipped pprof binaries into a JSON‑L trace format, you can augment your repository's call graph with real execution paths observed in production.

## What are eBPF Production Profiles?

**eBPF continuous profilers** such as [Parca](https://parca.dev) and [Pyroscope](https://pyroscope.io) use kernel‑level instrumentation to sample CPU stacks with minimal overhead. These tools export data in the standard **pprof** binary format (protocol buffer‑encoded, gzip‑compressed), which contains:

- Sample timestamps and values
- Stack frame addresses and symbol names
- Mapping information for shared libraries

Code‑Graph‑RAG treats these profiles as **dynamic traces**—evidence of actual function invocations that may differ from static call‑graph predictions due to polymorphism, callbacks, or runtime plugin loading.

## The Three‑Step Integration Workflow

### 1. Export a pprof Profile from Your Profiler

Both Parca and Pyroscope expose HTTP endpoints for profile retrieval.

**Parca endpoint pattern:**

```bash
curl -L "https://parca.example.com/api/v1/profile?query=my-service&from=now-5m&to=now&format=pprof" \
     -o prod_profile.pb.gz

```

**Pyroscope endpoint pattern:**

```bash
curl -L "http://pyroscope-server:4040/render?query=my-service.cpu&from=now-5m&to=now&format=pprof" \
     -o prod_profile.pb.gz

```

The resulting file is a gzipped protobuf message conforming to the [pprof profile schema](https://github.com/google/pprof/blob/main/proto/profile.proto).

### 2. Convert the Profile with `cgr trace convert`

Code‑Graph‑RAG's CLI provides the `convert` subcommand under the `trace` namespace. According to the implementation in [[`codebase_rag/trace/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/cli.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/cli.py), this entry point delegates to `convert_ebpf_pprof` in [[`codebase_rag/trace/ebpf_pprof.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/ebpf_pprof.py)](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/ebpf_pprof.py).

**CLI conversion:**

```bash
cgr trace convert prod_profile.pb.gz \
    --format ebpf \
    --repo-path /path/to/your/repo \
    --output dynamic_trace.jsonl

```

**Parameter reference:**

| Flag | Description |
|------|-------------|
| `profile` (positional) | Path to the gzipped pprof file |
| `--format ebpf` | Required to trigger the eBPF‑specific parser |
| `--repo-path` | Root of the repository for path normalization |
| `--output` | Destination for the JSON‑L trace file |

### 3. Merge the Dynamic Trace with Your Static Graph

The converted trace contains line‑delimited JSON objects representing observed call stacks. Load these into your existing graph with:

```bash
cgr trace load \
    --graph graph.json \
    --trace dynamic_trace.jsonl \
    --output merged_graph.json

```

## How the Conversion Works Internally

The `convert_ebpf_pprof` function in [`codebase_rag/trace/ebpf_pprof.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/ebpf_pprof.py) performs these operations:

1. **Parse the protobuf** using the vendored [`schema_pb2.py`](https://github.com/vitali87/code-graph-rag/blob/main/schema_pb2.py) definitions
2. **Iterate samples** and reconstruct call stacks from `sample.location_id` chains
3. **Normalize paths** against the provided `--repo-path` to enable source linking
4. **Emit JSON‑L records** with schema: `{"caller": "funcA", "callee": "funcB", "source": "ebpf", "weight": N}`

Each record includes a **weight** derived from sample values (CPU time or sample count), allowing downstream analysis to prioritize hot paths.

## Python API for Programmatic Integration

For custom pipelines, use the Python API directly:

```python
from codebase_rag.trace.ebpf_pprof import convert_ebpf_pprof
from codebase_rag.graph_loader import load_graph, apply_trace
from pathlib import Path

# Step 1: Convert eBPF profile to trace format

profile_path = Path("prod_profile.pb.gz")
repo_root = Path("/path/to/your/repo")
trace_path = Path("dynamic_trace.jsonl")

sample_count = convert_ebpf_pprof(
    profile_path,
    repo_path=repo_root,
    output_path=trace_path,
    format="ebpf",
)
print(f"Converted {sample_count} samples to {trace_path}")

# Step 2: Load and merge with static graph

graph = load_graph("graph.json")
graph = apply_trace(graph, trace_path, trace_source="ebpf")
graph.save("merged_graph.json")

```

The `apply_trace` function distinguishes dynamic edges by tagging them with `source: "ebpf"`, enabling filtered queries like: *"Show me all static calls that were never observed in production."*

## Handling Symbolization and Path Mapping

Production eBPF profiles may contain **raw addresses** rather than symbol names if the profiler lacks access to debug symbols. The conversion pipeline assumes server‑side symbolization (Parca/Pyroscope's default) but provides hooks for:

- **Offline symbol resolution** via `addr2line` or `llvm-symbolizer`
- **Path remapping** for reproducible builds where source paths differ between compilation and analysis environments

Configure remapping via environment variable or CLI flag as documented in [[`docs/guide/dynamic-tracing.md`](https://github.com/vitali87/code-graph-rag/blob/main/docs/guide/dynamic-tracing.md)](https://github.com/vitali87/code-graph-rag/blob/main/docs/guide/dynamic-tracing.md):

```bash
CGR_PATH_MAP="/build/path:/repo/path" cgr trace convert ...

```

## CI/CD Integration Example

Automate profile ingestion in your deployment pipeline:

```yaml

# .github/workflows/profile-ingest.yml

name: Ingest Production Profiles

on:
  schedule:
    - cron: '0 */6 * * *'  # Every 6 hours

jobs:
  ingest:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      
      - name: Install code-graph-rag
        run: pip install code-graph-rag
      
      - name: Fetch latest Parca profile
        run: |
          curl -L "${{ secrets.PARCA_URL }}/api/v1/profile?query=prod-api&from=now-6h&format=pprof" \
               -o profile.pb.gz
      
      - name: Convert and merge
        run: |
          cgr trace convert profile.pb.gz --format ebpf --repo-path . --output trace.jsonl
          cgr trace load --graph base_graph.json --trace trace.jsonl --output graph.json
      
      - name: Upload enriched graph
        uses: actions/upload-artifact@v4
        with:
          name: enriched-graph
          path: graph.json

```

## Summary

- **Export** pprof profiles from Parca, Pyroscope, or OpenTelemetry eBPF profilers via their HTTP APIs
- **Convert** using `cgr trace convert --format ebpf`, implemented in [`codebase_rag/trace/ebpf_pprof.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/ebpf_pprof.py)
- **Merge** the resulting JSON‑L trace with your static graph using `cgr trace load` or the Python API
- Distinguish **dynamic eBPF edges** from static analysis via the `source` metadata field
- Automate ingestion in CI pipelines for continuously updated production‑informed call graphs

## Frequently Asked Questions

### What eBPF profilers are officially supported by Code‑Graph‑RAG?

The [`ebpf_pprof.py`](https://github.com/vitali87/code-graph-rag/blob/main/ebpf_pprof.py) converter handles any profiler emitting **standard pprof format**, including Parca, Pyroscope, Grafana Cloud Profiles, and the OpenTelemetry eBPF profiler. The key requirement is gzip‑compressed protobuf output conforming to the pprof schema.

### How do I handle profiles with unsymbolized addresses?

Ensure your profiler performs **server‑side symbolization**—both Parca and Pyroscope support uploading debug info or fetching symbols from symbol servers. For offline processing, pre‑process profiles with `pprof` CLI tools to resolve addresses before conversion.

### Can I weight edges by CPU time versus sample count?

Yes. The converter reads the `sample.value` array and uses the **default sample type** (typically CPU nanoseconds or samples) as the edge weight. The protobuf `sample_type` field determines which value is extracted; this matches pprof's semantic conventions.

### What's the performance impact of converting large production profiles?

Conversion is **streaming and memory‑bounded**: [`ebpf_pprof.py`](https://github.com/vitali87/code-graph-rag/blob/main/ebpf_pprof.py) processes samples iteratively without loading the entire profile into RAM. Typical throughput exceeds 100,000 samples/second on modern hardware. For multi‑gigabyte profiles, consider the `--samples-limit` flag to cap ingestion.