How magic-trace perf_dlfilter Resolves Symbols: The Dynamic Linker Filter Mechanism
The perfdlfilter component in magic-trace acts as a dynamic linker filter that intercepts raw instruction pointers from Linux perf events, walks the process link-map via dl_iterate_phdr to identify ELF objects, resolves symbols through the OCaml Symbol module, and caches mappings to efficiently translate program counters into human-readable function names and source locations.
When tracing performance with janestreet/magic-trace, raw instruction pointers captured by the Linux kernel require translation into meaningful symbols. The perfdlfilter mechanism bridges this gap by operating as a specialized filter within the magic-trace pipeline. It processes perf event streams in real-time, resolving addresses against ELF metadata and debug information without requiring process restarts.
What Is the magic-trace perf_dlfilter?
perfdlfilter is the dynamic linker filter component inside magic-trace responsible for converting opaque program counter (PC) values into actionable symbol information. Unlike traditional offline symbolization, this mechanism operates as a callback registered with the perf event reader, enabling live resolution as samples arrive from the kernel's perf subsystem.
The filter integrates tightly with the runtime loader's interface, allowing magic-trace to query loaded shared objects and their symbol tables on-the-fly. This design supports JIT-generated code and dynamically loaded libraries without prior knowledge of the process memory layout.
How the Symbol Resolution Pipeline Works
The perfdlfilter mechanism processes each sample through a six-stage pipeline that transforms raw PERF_RECORD_SAMPLE data into enriched trace events.
Stage 1: Intercepting Raw Perf Events
The Linux kernel delivers raw samples via the perf subsystem, each containing a program counter value representing the executing instruction. These events arrive as PERF_RECORD_SAMPLE records containing the PC and the PID/TID of the emitting thread.
In src/perf_decode.ml, the raw byte stream undergoes initial parsing before reaching the symbol resolution layer.
Stage 2: The perf_dlfilter Callback
The C implementation in src/perf_dlfilter.c registers a callback with the perf-event reader. When a sample arrives, this callback receives the raw PC value along with the process and thread identifiers.
The header file src/perf_dlfilter.h exposes these C functions to the OCaml runtime through FFI bindings, allowing the OCaml Perf_event module to set the callback:
(* Registration via OCaml binding *)
let register () =
Perf_event.set_dlfilter_callback ~cb:Perf_dlfilter.callback
Stage 3: ELF Object Lookup via dl_iterate_phdr
To determine which shared object or executable contains the address, the callback uses the dl_iterate_phdr API to walk the ELF program headers of all loaded objects in the target process. The implementation iterates through the link-map until finding the object whose load address range contains the specific PC value.
This operation occurs in src/perf_dlfilter.c, which handles the low-level C structures and memory mappings.
Stage 4: Symbol Resolution with the Symbol Module
Once the owning ELF object is identified, magic-trace queries the object's symbol table. The Symbol module in src/symbol.ml performs this lookup, searching .symtab and .dynsym sections for the nearest symbol below the target address.
When DWARF-encoded debugging data is present, the module can additionally resolve source file names and line numbers. The interface defined in src/symbol.mli provides the resolve function that maps object offsets to symbol names.
let callback ~pid ~tid ~pc =
match Symbol_cache.find (pid, pc) with
| Some sym -> sym
| None ->
let obj = Dl.iterate_phdr pid pc in (* locate ELF object *)
let sym = Symbol.resolve ~obj ~pc in (* find nearest symbol *)
Symbol_cache.add (pid, pc) sym;
sym
Stage 5: Address-to-Symbol Caching
To minimize overhead during high-frequency sampling, resolved mappings are stored in a hash table keyed by (pid, pc) tuples. This caching strategy ensures that expensive ELF parsing and symbol table scans occur only once per unique address.
Subsequent samples hitting the same instruction pointer bypass the link-map walk and symbol lookup, retrieving the cached result instantly.
Stage 6: Emitting Resolved Symbols to the Trace Pipeline
The final stage converts the resolved symbol information into a structured trace event. The mapping—containing the object name, offset within the object, and symbol name—is wrapped in a Trace_event.Symbol record as defined in src/trace.ml.
Downstream components such as the call-stack decoder in src/trace_filter.ml consume these events to render human-readable stack traces:
let emit_sample ~pid ~tid ~pc =
let sym = Perf_dlfilter.resolve ~pid ~pc in
Trace_event.emit (Sample { pid; tid; pc; symbol = sym })
Key Implementation Files in janestreet/magic-trace
The perfdlfilter mechanism spans multiple source files across C and OCaml:
src/perf_dlfilter.c— C implementation of the dlfilter callback. Handlesdl_iterate_phdriteration and link-map traversal.src/perf_dlfilter.h— Header file exposing C functions to OCaml via FFI.src/symbol.ml— Core logic for reading ELF symbol tables (.symtab,.dynsym) and DWARF debug information.src/symbol.mli— Interface signature for the symbol resolution module.src/trace.ml— Defines trace event types including theSamplerecord that carries resolved symbols.src/perf_decode.ml— Decodes raw perf events and orchestrates dlfilter invocation during sample processing.src/trace_filter.ml— Higher-level filter pipeline that integrates dlfilter output with other trace processing stages.
Practical Benefits: JIT Support and Performance
The dynamic nature of perfdlfilter provides specific advantages for modern tracing scenarios:
JIT-Generated Code Resolution — When JIT libraries (such as JavaScript engines or LLVM-based compilers) allocate new code segments, dl_iterate_phdr immediately sees the new mappings. Subsequent samples from these regions resolve correctly without restarting the trace.
On-the-Fly Debug Info Loading — If debug information files appear after process startup, the filter re-queries the ELF file during resolution. This allows developers to attach debugging symbols to running processes without interruption.
Minimal Overhead Through Caching — The (pid, pc) keyed cache eliminates redundant ELF I/O and symbol table searches. According to the janestreet/magic-trace source code, this optimization is crucial for handling high-frequency sampling without dropping events.
Summary
perfdlfilteracts as a dynamic linker filter within magic-trace, intercepting raw perf events to resolve instruction pointers into symbols.- The mechanism uses
dl_iterate_phdrinsrc/perf_dlfilter.cto map addresses to ELF objects, then queriessrc/symbol.mlfor specific function names and source locations. - A hash table cache keyed by process ID and program counter prevents repeated expensive lookups, enabling high-frequency tracing.
- The resolved symbols flow through
src/trace.mlandsrc/trace_filter.mlto produce human-readable stack traces. - The design supports JIT-generated code and dynamically loaded debug information without requiring process restarts.
Frequently Asked Questions
What is the perf_dlfilter mechanism in magic-trace?
The perfdlfilter mechanism is a dynamic linker filter that intercepts raw instruction pointers from Linux perf samples. It resolves these addresses into human-readable symbols by querying the process's ELF metadata and runtime link-map. This occurs live during tracing, allowing magic-trace to display function names and source locations immediately as samples are captured.
How does magic-trace resolve symbols for JIT-compiled code?
Magic-trace resolves JIT-compiled code by leveraging the dl_iterate_phdr API to detect newly loaded code segments dynamically. When a JIT allocates executable memory and updates the process link-map, subsequent perf samples from that address range trigger a fresh lookup in src/perf_dlfilter.c. The filter discovers the new mapping and resolves symbols against the JIT's exported symbol table or debug information.
Where is the symbol resolution logic implemented in the magic-trace repository?
The primary symbol resolution logic resides in src/symbol.ml, which reads ELF symbol tables (.symtab and .dynsym) and DWARF debug data. The C-based link-map traversal lives in src/perf_dlfilter.c, while the integration with the trace pipeline occurs in src/perf_decode.ml and src/trace_filter.ml. The interface between C and OCaml is defined in src/perf_dlfilter.h.
Why does magic-trace cache symbol resolution results?
Magic-trace caches results in a hash table keyed by (pid, pc) to eliminate redundant ELF parsing and symbol table scans. Since high-frequency sampling generates many events for the same addresses, caching ensures that the expensive dl_iterate_phdr walk and symbol lookup occur only once per unique instruction pointer. This optimization is essential for maintaining low overhead during intensive tracing sessions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →