Dwarf vs LBR vs Frame Pointers: Comparing Magic-Trace Callgraph Modes

Magic-trace supports three distinct callgraph modes—DWARF, Last Branch Record (LBR), and Frame Pointers—that trade off between portability, hardware requirements, and runtime overhead when reconstructing stack traces from Linux perf samples.

Magic-trace, Jane Street's open-source tracer built on the Linux perf subsystem, reconstructs execution call stacks by configuring specific magic-trace callgraph modes that determine how the unwinder interprets binary and CPU state. These modes define whether the tool relies on debug metadata, hardware branch buffers, or register-based frame chains to build call trees. Understanding the architectural differences between DWARF, Last Branch Record, and Frame Pointers is crucial for selecting the optimal strategy for your hardware capabilities and compilation environment.

How Callgraph Modes Work

In src/callgraph_mode.ml, magic-trace defines the callgraph_mode type enumerating three variants: Last_branch_record (with a stitched boolean flag), Frame_pointers, and Dwarf. When recording, the tool translates these internal representations into corresponding perf command-line arguments (--call-graph lbr, --call-graph fp, or --call-graph dwarf) to control how the Linux kernel collects stack traces.

DWARF Mode: Debug Information Unwinding

DWARF mode relies on DWARF debugging information embedded in binaries to unwind the stack. It reads line tables and function ranges from debug sections to map instruction pointers back to source locations and call frames.

This mode works on any architecture and requires no special CPU features or compilation flags. Because it depends solely on debug metadata, DWARF operates on any binary that includes debugging information, eliminating the need to recompile target programs.

The trade-off is performance. DWARF decoding introduces higher runtime overhead than hardware-assisted methods, produces larger perf.data files, and requires more CPU time to resolve each sampled instruction pointer against debug entries.

Last Branch Record (LBR) Mode: Hardware-Assisted Tracing

Last Branch Record (LBR) leverages Intel CPU hardware buffers that store recent branch target addresses. When available, magic-trace uses this capability to reconstruct call stacks with minimal overhead by reading the CPU's internal branch history.

This mode requires Intel processors that expose the last_branch_record capability, which magic-trace detects via Perf_capabilities.do_intersect in src/perf_tool_backend.ml. When LBR is available, it becomes the default mode due to its low overhead and fast reconstruction—perf already records branch targets without expensive software unwinding.

The limitation is buffer depth. The LBR hardware stores only a fixed number of recent branches, limiting call-stack depth. Additionally, this mode is Intel-specific and unavailable on other architectures.

LBR Stitching

Magic-trace supports two LBR variants controlled by the stitched flag. Stitched mode (--call-graph lbr) chains consecutive LBR entries to reconstruct longer call stacks than a single hardware buffer would allow. Non-stitched mode (lbr-no-stitch) provides raw LBR data without this chaining, useful for low-level branch analysis.

When stitching is enabled, the tool attempts to combine entries into a pseudo call-stack; disabling it yields raw branch sequences.

Frame Pointers Mode: Register-Based Unwinding

Frame Pointers mode uses the frame-pointer register (%rbp on x86_64) to walk the stack linearly. This requires the target binary to be compiled with -fno-omit-frame-pointers (or equivalent), ensuring the compiler maintains the frame-pointer chain.

This mode works on any architecture but requires specific compilation flags. It provides a middle ground: it requires no special CPU features like LBR and no debug metadata like DWARF, but demands build-system cooperation to retain frame pointers.

If the binary lacks frame pointers, the unwinder fails to reconstruct the stack. Performance characteristics typically fall between LBR (fastest) and DWARF (slowest), though still requiring software-based traversal.

Automatic Mode Selection and Fallbacks

In src/perf_tool_backend.ml, magic-trace implements automatic fallback logic when you omit the -callgraph-mode flag. The tool queries perf capabilities to check for last_branch_record support.

If the CPU supports LBR, magic-trace defaults to Last_branch_record with stitching enabled. If LBR is unavailable—common on non-Intel hardware or virtualized environments—the tool automatically falls back to DWARF mode and emits a warning explaining the trade-off.

This capability-aware selection ensures tracing works across diverse environments while optimizing for low overhead when hardware support is present.

Command-Line Usage Examples

Control the magic-trace callgraph mode explicitly using the -callgraph-mode flag. Here are practical examples for each mode:

Use DWARF for maximum portability on any architecture or binary:

magic-trace record -callgraph-mode dwarf -o output.perf ./my_program

Enable LBR with stitching (default on supported Intel CPUs):

magic-trace record -callgraph-mode lbr -o output.perf ./my_program

Capture raw LBR data without stitching:

magic-trace record -callgraph-mode lbr-no-stitch -o output.perf ./my_program

Require frame pointers (ensure your binary was compiled with -fno-omit-frame-pointers):

magic-trace record -callgraph-mode fp -o output.perf ./my_program

When the -callgraph-mode flag is omitted, magic-trace selects the best available mode based on hardware capabilities and prints an informational message describing the active configuration.

Summary

  • DWARF mode uses debug information to unwind stacks, works on any architecture without recompilation, but incurs higher overhead and larger trace files.
  • Last Branch Record (LBR) mode leverages Intel CPU hardware buffers for minimal overhead and fast decoding, but requires compatible Intel processors and limits stack depth to the hardware buffer size.
  • Frame Pointers mode relies on compiler-generated frame chains, requires binaries built with -fno-omit-frame-pointers, and offers compatibility without debug metadata or special CPU features.
  • Magic-trace automatically selects LBR when available, falling back to DWARF on unsupported hardware, with explicit control available via the -callgraph-mode flag in src/callgraph_mode.ml.

Frequently Asked Questions

Which magic-trace callgraph mode provides the lowest overhead?

Last Branch Record (LBR) mode provides the lowest overhead because it reads branch targets directly from Intel CPU hardware buffers without software unwinding. This eliminates the debug-info parsing required by DWARF and the pointer-chasing needed for Frame Pointers, resulting in smaller perf.data files and reduced CPU utilization during tracing.

Why does magic-trace default to DWARF instead of LBR on my system?

Magic-trace defaults to DWARF when the CPU lacks the last_branch_record capability, which is exclusive to modern Intel processors. In src/perf_tool_backend.ml, the tool checks Perf_capabilities.do_intersect capabilities last_branch_record and automatically falls back to DWARF with a warning when LBR is unavailable, such as on AMD processors, older Intel chips, or virtualized environments without hardware pass-through.

Do I need to recompile my program to use Frame Pointers mode?

Yes. Frame Pointers mode requires the target binary to be compiled with -fno-omit-frame-pointers (GCC/Clang) or equivalent flags. Without this flag, compilers optimize away the frame pointer register (%rbp), breaking the linked-list chain that the unwinder traverses. Unlike DWARF, which works on existing binaries with debug symbols, Frame Pointers mode strictly requires build-system modification and recompilation.

What is the difference between lbr and lbr-no-stitch in magic-trace?

The stitched variant (-callgraph-mode lbr) chains consecutive LBR hardware buffer entries to synthesize longer call stacks than the physical buffer allows, while non-stitched (-callgraph-mode lbr-no-stitch) outputs raw LBR records without this chaining logic. Use the stitched mode for typical call-stack reconstruction and non-stitched mode when you need raw branch target data for low-level analysis of control-flow transfers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →