LLVM Performance Profiling with `perf` and Intel VTune: A Complete Guide
LLVM provides two primary mechanisms for performance profiling: the llvm-profgen tool converts Linux perf recordings into optimizable sample profiles, while the VTune support plugin registers JIT-compiled code with Intel VTune Amplifier for hotspot analysis.
The LLVM project (llvm/llvm-project) offers deep integration with industry-standard profilers to drive profile-guided optimizations (PGO) and analyze runtime behavior. Whether you are optimizing Ahead-of-Time (AOT) compiled binaries using hardware sampling or debugging Just-in-Time (JIT) generated machine code, LLVM exposes specific command-line options and compiler flags to bridge these tools with its optimization pipeline.
Profiling with Linux perf
LLVM leverages the Linux perf subsystem through the llvm-profgen utility, which translates raw perf recordings into the LLVM Sample Profile format (.prof). This workflow enables accurate branch-stack analysis and hot-path identification without instrumentation overhead.
Recording Performance Data
To collect samples suitable for LLVM consumption, record hardware events using perf with branch-stack capturing enabled. The -b flag ensures Last Branch Record (LBR) or Branch Record Buffer Extensions (BRBE) stacks are captured, which llvm-profgen uses to reconstruct call contexts.
# Record raw perf data with branch stacks
perf record -e cycles -b -o app.perfdata -- ./myapp
# Generate human-readable script for troubleshooting
perf script -i app.perfdata > app.perfscript
Converting perf Data to LLVM Sample Profiles
The conversion happens in llvm/tools/llvm-profgen/llvm-profgen.cpp, which enforces mutually exclusive input options (--perfscript, --perfdata, --etm) as validated in lines 32-38 and the validateCommandLine function (lines 105-119).
Using raw binary data:
llvm-profgen --perfdata app.perfdata \
--binary ./myapp \
-o app.prof
Using a perf script:
llvm-profgen --perfscript app.perfscript \
--binary ./myapp \
-o app.prof
Under the hood, PerfReader.cpp parses the raw perf files to extract events, while ProfileGenerator.cpp builds the sample profile by aggregating these events into the format expected by LLVM's optimizer.
Consuming Profiles in the Optimizer
Once generated, profiles can be merged across multiple runs using llvm-profdata merge, then fed into the compiler via -fprofile-use:
# Merge multiple profile runs
llvm-profdata merge -output=merged.prof run1.prof run2.prof
# Compile with profile-guided optimization
clang -O2 -fprofile-use=merged.prof -c myapp.c
Intel VTune Integration for JIT Code
For JIT compilation scenarios, LLVM provides the VTune Support Plugin, which registers dynamically generated code objects with Intel VTune Amplifier. This allows VTune to attribute hardware samples to specific JIT-compiled functions and display source correspondence.
How VTune Support Works
The integration centers on llvm/lib/ExecutionEngine/Orc/Debugging/VTuneSupportPlugin.cpp. When JIT code is emitted, the VTuneSupportPlugin::getMethodBatch method (lines 26-83) constructs a VTuneMethodBatch containing the function's load address, size, and optional DWARF line tables. This batch is forwarded to the VTune runtime via wrapper functions defined in JITLoaderVTune.cpp (llvm_orc_registerVTuneImpl and llvm_orc_unregisterVTuneImpl).
Enabling VTune in JIT Applications
Activation requires building LLVM with VTune support enabled and using the appropriate runtime flags.
Build configuration:
cmake -DLLVM_ENABLE_VTUNE=ON <llvm-source-dir>
Runtime usage with llvm-jitlink:
The tool llvm-jitlink exposes the --vtune-support flag (defined in llvm/tools/llvm-jitlink/llvm-jitlink.cpp at line 225) to activate the plugin:
llvm-jitlink --vtune-support --binary ./jit_program
Profiling workflow:
# Run under VTune to collect hotspots
vtune -collect hotspots -- ./jit_program
VTune will then display JIT regions with their associated source locations and sample counts, enabling precise performance analysis of dynamically generated code.
Additional Profiling Capabilities
Polly Performance Monitoring
The Polly polyhedral optimizer can emit perf-compatible instrumentation for loop-nest analysis. Enable the -polly-codegen-perf-monitoring pass (implemented in llvm/polly/lib/CodeGen/CodeGeneration.cpp at line 58) to generate counters for trip counts and cycle measurements:
opt -load libPolly.so -polly-codegen-perf-monitoring -S < input.ll > output.ll
ARM ETM and Synthetic Counters
For ARM architecture trace macrocell (ETM) data, llvm-profgen accepts the --etm flag (lines 91-98 in llvm-profgen.cpp) to process embedded trace streams. Additionally, the llvm-exegesis benchmarking tool supports --use-dummy-perf-counters (defined in llvm/tools/llvm-exegesis/llvm-exegesis.cpp at line 139) to run performance measurements without accessing physical hardware counters.
Summary
llvm-profgenconverts Linuxperfrecordings (raw data or scripts) into LLVM Sample Profiles using--perfdataor--perfscript, with source files located inllvm/tools/llvm-profgen/.- Intel VTune integration for JIT code requires building with
-DLLVM_ENABLE_VTUNE=ONand using the VTune Support Plugin (VTuneSupportPlugin.cpp) to register code objects viaJITLoaderVTune.cpp. - Profile consumption involves
llvm-profdata mergeand the-fprofile-usecompiler flag to drive optimizations based on collected samples. - Extended options include ARM ETM trace support (
--etm), Polly performance monitoring (-polly-codegen-perf-monitoring), and synthetic counters forllvm-exegesis.
Frequently Asked Questions
How do I convert an existing perf.data file into a format LLVM can use?
Use the llvm-profgen tool with the --perfdata flag to convert raw perf recordings into LLVM Sample Profile format. Specify the binary with --binary to ensure proper symbolization, then output to a .prof file that can be consumed by clang -fprofile-use.
What is the difference between using --perfscript and --perfdata in llvm-profgen?
The --perfdata option accepts raw binary perf recordings (typically created with perf record), while --perfscript expects the textual output generated by perf script. Both produce identical .prof files, but raw data is smaller and faster to process, whereas scripts are human-readable and useful for debugging.
Why can't I see my JIT-compiled functions in Intel VTune?
JIT functions require explicit registration with the VTune runtime. Ensure LLVM was built with -DLLVM_ENABLE_VTUNE=ON, and that your JIT execution engine loads the VTune Support Plugin. If using llvm-jitlink, add the --vtune-support flag to enable the registration batch that communicates function addresses and debug information to VTune.
Does LLVM support profiling on ARM architectures beyond standard perf?
Yes, llvm-profgen supports the --etm flag for processing ARM Embedded Trace Macrocell (ETM) data files, enabling profiling on ARM targets where hardware tracing is available in addition to standard statistical sampling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →