# LLVM Performance Profiling with `perf` and Intel VTune: A Complete Guide

> Master LLVM performance profiling with Linux perf and Intel VTune. Learn to convert perf recordings with llvm-profgen and analyze JIT-compiled code with VTune for optimal performance.

- Repository: [LLVM/llvm-project](https://github.com/llvm/llvm-project)
- Tags: deep-dive
- Published: 2026-09-11

---

**LLVM provides two primary mechanisms for performance profiling: the `llvm-profgen` tool converts Linux `perf` recordings into optimizable sample profiles, while the VTune support plugin registers JIT-compiled code with Intel VTune Amplifier for hotspot analysis.**

The LLVM project (`llvm/llvm-project`) offers deep integration with industry-standard profilers to drive profile-guided optimizations (PGO) and analyze runtime behavior. Whether you are optimizing Ahead-of-Time (AOT) compiled binaries using hardware sampling or debugging Just-in-Time (JIT) generated machine code, LLVM exposes specific command-line options and compiler flags to bridge these tools with its optimization pipeline.

## Profiling with Linux `perf`

LLVM leverages the Linux `perf` subsystem through the **`llvm-profgen`** utility, which translates raw perf recordings into the LLVM Sample Profile format (`.prof`). This workflow enables accurate branch-stack analysis and hot-path identification without instrumentation overhead.

### Recording Performance Data

To collect samples suitable for LLVM consumption, record hardware events using `perf` with branch-stack capturing enabled. The `-b` flag ensures Last Branch Record (LBR) or Branch Record Buffer Extensions (BRBE) stacks are captured, which `llvm-profgen` uses to reconstruct call contexts.

```bash

# Record raw perf data with branch stacks

perf record -e cycles -b -o app.perfdata -- ./myapp

# Generate human-readable script for troubleshooting

perf script -i app.perfdata > app.perfscript

```

### Converting perf Data to LLVM Sample Profiles

The conversion happens in [`llvm/tools/llvm-profgen/llvm-profgen.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm/tools/llvm-profgen/llvm-profgen.cpp), which enforces mutually exclusive input options (`--perfscript`, `--perfdata`, `--etm`) as validated in lines 32-38 and the `validateCommandLine` function (lines 105-119).

**Using raw binary data:**

```bash
llvm-profgen --perfdata app.perfdata \
             --binary ./myapp \
             -o app.prof

```

**Using a perf script:**

```bash
llvm-profgen --perfscript app.perfscript \
             --binary ./myapp \
             -o app.prof

```

Under the hood, [`PerfReader.cpp`](https://github.com/llvm/llvm-project/blob/main/PerfReader.cpp) parses the raw perf files to extract events, while [`ProfileGenerator.cpp`](https://github.com/llvm/llvm-project/blob/main/ProfileGenerator.cpp) builds the sample profile by aggregating these events into the format expected by LLVM's optimizer.

### Consuming Profiles in the Optimizer

Once generated, profiles can be merged across multiple runs using `llvm-profdata merge`, then fed into the compiler via `-fprofile-use`:

```bash

# Merge multiple profile runs

llvm-profdata merge -output=merged.prof run1.prof run2.prof

# Compile with profile-guided optimization

clang -O2 -fprofile-use=merged.prof -c myapp.c

```

## Intel VTune Integration for JIT Code

For JIT compilation scenarios, LLVM provides the **VTune Support Plugin**, which registers dynamically generated code objects with Intel VTune Amplifier. This allows VTune to attribute hardware samples to specific JIT-compiled functions and display source correspondence.

### How VTune Support Works

The integration centers on [`llvm/lib/ExecutionEngine/Orc/Debugging/VTuneSupportPlugin.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm/lib/ExecutionEngine/Orc/Debugging/VTuneSupportPlugin.cpp). When JIT code is emitted, the `VTuneSupportPlugin::getMethodBatch` method (lines 26-83) constructs a `VTuneMethodBatch` containing the function's load address, size, and optional DWARF line tables. This batch is forwarded to the VTune runtime via wrapper functions defined in [`JITLoaderVTune.cpp`](https://github.com/llvm/llvm-project/blob/main/JITLoaderVTune.cpp) (`llvm_orc_registerVTuneImpl` and `llvm_orc_unregisterVTuneImpl`).

### Enabling VTune in JIT Applications

Activation requires building LLVM with VTune support enabled and using the appropriate runtime flags.

**Build configuration:**

```bash
cmake -DLLVM_ENABLE_VTUNE=ON <llvm-source-dir>

```

**Runtime usage with llvm-jitlink:**

The tool `llvm-jitlink` exposes the `--vtune-support` flag (defined in [`llvm/tools/llvm-jitlink/llvm-jitlink.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm/tools/llvm-jitlink/llvm-jitlink.cpp) at line 225) to activate the plugin:

```bash
llvm-jitlink --vtune-support --binary ./jit_program

```

**Profiling workflow:**

```bash

# Run under VTune to collect hotspots

vtune -collect hotspots -- ./jit_program

```

VTune will then display JIT regions with their associated source locations and sample counts, enabling precise performance analysis of dynamically generated code.

## Additional Profiling Capabilities

### Polly Performance Monitoring

The Polly polyhedral optimizer can emit `perf`-compatible instrumentation for loop-nest analysis. Enable the `-polly-codegen-perf-monitoring` pass (implemented in [`llvm/polly/lib/CodeGen/CodeGeneration.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm/polly/lib/CodeGen/CodeGeneration.cpp) at line 58) to generate counters for trip counts and cycle measurements:

```bash
opt -load libPolly.so -polly-codegen-perf-monitoring -S < input.ll > output.ll

```

### ARM ETM and Synthetic Counters

For ARM architecture trace macrocell (ETM) data, `llvm-profgen` accepts the `--etm` flag (lines 91-98 in [`llvm-profgen.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm-profgen.cpp)) to process embedded trace streams. Additionally, the `llvm-exegesis` benchmarking tool supports `--use-dummy-perf-counters` (defined in [`llvm/tools/llvm-exegesis/llvm-exegesis.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm/tools/llvm-exegesis/llvm-exegesis.cpp) at line 139) to run performance measurements without accessing physical hardware counters.

## Summary

- **`llvm-profgen`** converts Linux `perf` recordings (raw data or scripts) into LLVM Sample Profiles using `--perfdata` or `--perfscript`, with source files located in `llvm/tools/llvm-profgen/`.
- **Intel VTune integration** for JIT code requires building with `-DLLVM_ENABLE_VTUNE=ON` and using the VTune Support Plugin ([`VTuneSupportPlugin.cpp`](https://github.com/llvm/llvm-project/blob/main/VTuneSupportPlugin.cpp)) to register code objects via [`JITLoaderVTune.cpp`](https://github.com/llvm/llvm-project/blob/main/JITLoaderVTune.cpp).
- **Profile consumption** involves `llvm-profdata merge` and the `-fprofile-use` compiler flag to drive optimizations based on collected samples.
- **Extended options** include ARM ETM trace support (`--etm`), Polly performance monitoring (`-polly-codegen-perf-monitoring`), and synthetic counters for `llvm-exegesis`.

## Frequently Asked Questions

### How do I convert an existing perf.data file into a format LLVM can use?

Use the `llvm-profgen` tool with the `--perfdata` flag to convert raw `perf` recordings into LLVM Sample Profile format. Specify the binary with `--binary` to ensure proper symbolization, then output to a `.prof` file that can be consumed by `clang -fprofile-use`.

### What is the difference between using `--perfscript` and `--perfdata` in llvm-profgen?

The `--perfdata` option accepts raw binary `perf` recordings (typically created with `perf record`), while `--perfscript` expects the textual output generated by `perf script`. Both produce identical `.prof` files, but raw data is smaller and faster to process, whereas scripts are human-readable and useful for debugging.

### Why can't I see my JIT-compiled functions in Intel VTune?

JIT functions require explicit registration with the VTune runtime. Ensure LLVM was built with `-DLLVM_ENABLE_VTUNE=ON`, and that your JIT execution engine loads the VTune Support Plugin. If using `llvm-jitlink`, add the `--vtune-support` flag to enable the registration batch that communicates function addresses and debug information to VTune.

### Does LLVM support profiling on ARM architectures beyond standard perf?

Yes, `llvm-profgen` supports the `--etm` flag for processing ARM Embedded Trace Macrocell (ETM) data files, enabling profiling on ARM targets where hardware tracing is available in addition to standard statistical sampling.