How to Integrate Production eBPF Profiles (Parca, Pyroscope) with Code‑Graph‑RAG Tracing
TLDR: Code‑Graph‑RAG ingests pprof profiles from eBPF continuous profilers like Parca or Pyroscope via the cgr trace convert command, converting production call stacks into dynamic CALLS edges that merge with the static analysis graph.
The vitali87/code-graph-rag repository provides a dedicated conversion pipeline that bridges the gap between runtime eBPF profiling and static code visualization. By converting gzipped pprof binaries into a JSON‑L trace format, you can augment your repository's call graph with real execution paths observed in production.
What are eBPF Production Profiles?
eBPF continuous profilers such as Parca and Pyroscope use kernel‑level instrumentation to sample CPU stacks with minimal overhead. These tools export data in the standard pprof binary format (protocol buffer‑encoded, gzip‑compressed), which contains:
- Sample timestamps and values
- Stack frame addresses and symbol names
- Mapping information for shared libraries
Code‑Graph‑RAG treats these profiles as dynamic traces—evidence of actual function invocations that may differ from static call‑graph predictions due to polymorphism, callbacks, or runtime plugin loading.
The Three‑Step Integration Workflow
1. Export a pprof Profile from Your Profiler
Both Parca and Pyroscope expose HTTP endpoints for profile retrieval.
Parca endpoint pattern:
curl -L "https://parca.example.com/api/v1/profile?query=my-service&from=now-5m&to=now&format=pprof" \
-o prod_profile.pb.gz
Pyroscope endpoint pattern:
curl -L "http://pyroscope-server:4040/render?query=my-service.cpu&from=now-5m&to=now&format=pprof" \
-o prod_profile.pb.gz
The resulting file is a gzipped protobuf message conforming to the pprof profile schema.
2. Convert the Profile with cgr trace convert
Code‑Graph‑RAG's CLI provides the convert subcommand under the trace namespace. According to the implementation in [codebase_rag/trace/cli.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/cli.py), this entry point delegates to convert_ebpf_pprof in [codebase_rag/trace/ebpf_pprof.py](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/trace/ebpf_pprof.py).
CLI conversion:
cgr trace convert prod_profile.pb.gz \
--format ebpf \
--repo-path /path/to/your/repo \
--output dynamic_trace.jsonl
Parameter reference:
| Flag | Description |
|---|---|
profile (positional) |
Path to the gzipped pprof file |
--format ebpf |
Required to trigger the eBPF‑specific parser |
--repo-path |
Root of the repository for path normalization |
--output |
Destination for the JSON‑L trace file |
3. Merge the Dynamic Trace with Your Static Graph
The converted trace contains line‑delimited JSON objects representing observed call stacks. Load these into your existing graph with:
cgr trace load \
--graph graph.json \
--trace dynamic_trace.jsonl \
--output merged_graph.json
How the Conversion Works Internally
The convert_ebpf_pprof function in codebase_rag/trace/ebpf_pprof.py performs these operations:
- Parse the protobuf using the vendored
schema_pb2.pydefinitions - Iterate samples and reconstruct call stacks from
sample.location_idchains - Normalize paths against the provided
--repo-pathto enable source linking - Emit JSON‑L records with schema:
{"caller": "funcA", "callee": "funcB", "source": "ebpf", "weight": N}
Each record includes a weight derived from sample values (CPU time or sample count), allowing downstream analysis to prioritize hot paths.
Python API for Programmatic Integration
For custom pipelines, use the Python API directly:
from codebase_rag.trace.ebpf_pprof import convert_ebpf_pprof
from codebase_rag.graph_loader import load_graph, apply_trace
from pathlib import Path
# Step 1: Convert eBPF profile to trace format
profile_path = Path("prod_profile.pb.gz")
repo_root = Path("/path/to/your/repo")
trace_path = Path("dynamic_trace.jsonl")
sample_count = convert_ebpf_pprof(
profile_path,
repo_path=repo_root,
output_path=trace_path,
format="ebpf",
)
print(f"Converted {sample_count} samples to {trace_path}")
# Step 2: Load and merge with static graph
graph = load_graph("graph.json")
graph = apply_trace(graph, trace_path, trace_source="ebpf")
graph.save("merged_graph.json")
The apply_trace function distinguishes dynamic edges by tagging them with source: "ebpf", enabling filtered queries like: "Show me all static calls that were never observed in production."
Handling Symbolization and Path Mapping
Production eBPF profiles may contain raw addresses rather than symbol names if the profiler lacks access to debug symbols. The conversion pipeline assumes server‑side symbolization (Parca/Pyroscope's default) but provides hooks for:
- Offline symbol resolution via
addr2lineorllvm-symbolizer - Path remapping for reproducible builds where source paths differ between compilation and analysis environments
Configure remapping via environment variable or CLI flag as documented in [docs/guide/dynamic-tracing.md](https://github.com/vitali87/code-graph-rag/blob/main/docs/guide/dynamic-tracing.md):
CGR_PATH_MAP="/build/path:/repo/path" cgr trace convert ...
CI/CD Integration Example
Automate profile ingestion in your deployment pipeline:
# .github/workflows/profile-ingest.yml
name: Ingest Production Profiles
on:
schedule:
- cron: '0 */6 * * *' # Every 6 hours
jobs:
ingest:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install code-graph-rag
run: pip install code-graph-rag
- name: Fetch latest Parca profile
run: |
curl -L "${{ secrets.PARCA_URL }}/api/v1/profile?query=prod-api&from=now-6h&format=pprof" \
-o profile.pb.gz
- name: Convert and merge
run: |
cgr trace convert profile.pb.gz --format ebpf --repo-path . --output trace.jsonl
cgr trace load --graph base_graph.json --trace trace.jsonl --output graph.json
- name: Upload enriched graph
uses: actions/upload-artifact@v4
with:
name: enriched-graph
path: graph.json
Summary
- Export pprof profiles from Parca, Pyroscope, or OpenTelemetry eBPF profilers via their HTTP APIs
- Convert using
cgr trace convert --format ebpf, implemented incodebase_rag/trace/ebpf_pprof.py - Merge the resulting JSON‑L trace with your static graph using
cgr trace loador the Python API - Distinguish dynamic eBPF edges from static analysis via the
sourcemetadata field - Automate ingestion in CI pipelines for continuously updated production‑informed call graphs
Frequently Asked Questions
What eBPF profilers are officially supported by Code‑Graph‑RAG?
The ebpf_pprof.py converter handles any profiler emitting standard pprof format, including Parca, Pyroscope, Grafana Cloud Profiles, and the OpenTelemetry eBPF profiler. The key requirement is gzip‑compressed protobuf output conforming to the pprof schema.
How do I handle profiles with unsymbolized addresses?
Ensure your profiler performs server‑side symbolization—both Parca and Pyroscope support uploading debug info or fetching symbols from symbol servers. For offline processing, pre‑process profiles with pprof CLI tools to resolve addresses before conversion.
Can I weight edges by CPU time versus sample count?
Yes. The converter reads the sample.value array and uses the default sample type (typically CPU nanoseconds or samples) as the edge weight. The protobuf sample_type field determines which value is extracted; this matches pprof's semantic conventions.
What's the performance impact of converting large production profiles?
Conversion is streaming and memory‑bounded: ebpf_pprof.py processes samples iteratively without loading the entire profile into RAM. Typical throughput exceeds 100,000 samples/second on modern hardware. For multi‑gigabyte profiles, consider the --samples-limit flag to cap ingestion.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →