Performance Implications of Magic-Trace's Multi-Snapshot Mode

Enabling multi-snapshot mode in magic-trace adds approximately 8 microseconds of overhead per trigger hit and causes trace files to grow linearly with each snapshot, trading runtime performance and memory for continuous visibility into hot code paths.

Magic-trace is Jane Street's open-source execution tracing tool that captures high-resolution snapshots of running applications using Linux perf events. While the default single-snapshot mode stops tracing immediately after the first trigger fires, multi-snapshot mode (activated via the -multi-snapshot flag) keeps the trace collection active, generating a new snapshot on every trigger hit. This continuous collection introduces measurable performance penalties that developers must weigh against the benefit of detailed temporal data.

How Multi-Snapshot Mode Works

In single-snapshot mode (the default), magic-trace stops the perf ring-buffer and writes a single trace file the moment the trigger symbol is hit or the user sends an interrupt. In multi-snapshot mode, the tracer performs the following sequence on every trigger hit:

  • Stops the perf ring-buffer temporarily
  • Flushes pending events to disk
  • Writes snapshot metadata
  • Re-arms the breakpoint for the next hit

This cycle allows the application to continue running between snapshots, but each iteration incurs non-trivial overhead.

Per-Trigger Latency Overhead

The primary cost of multi-snapshot mode is approximately 8 microseconds (µs) of latency added to every trigger invocation. This overhead comes from kernel-level breakpoint handling and user-space buffer management required to extract each snapshot.

According to the source code in src/trace.ml (lines 538-447), the command-line flag documentation explicitly warns about this cost:

flag
  "-multi-snapshot"
  no_arg
  ~doc:
    "Take a snapshot every time the trigger is hit, instead of only the first \
     time. This flag has two caveats:\n\
     (1) There's an ~8us performance hit every time the trigger symbol is hit. If \
     snapshots trigger frequently, your application's performance may be \
     materially impacted.\n\
     (2) Each snapshot linearly increases the size of the trace file..."

For low-frequency triggers (e.g., application startup or rare error conditions), this 8 µs penalty is negligible. However, for hot-path functions called thousands of times per second, the accumulated latency can materially degrade application throughput.

Memory and Trace File Growth

Each snapshot captures a full copy of the current perf buffer plus associated metadata. In multi-snapshot mode, these snapshots accumulate linearly throughout the execution, directly increasing the final trace file size.

As noted in the flag documentation referenced above, this linear growth can exhaust available memory and may cause the trace viewer to crash when loading large files. The actual write operations are handled in src/trace_writer.ml, where each snapshot appends to the output stream without overwriting previous data.

Kernel Breakpoint Handling Costs

When multi-snapshot mode is active, the trigger breakpoint remains enabled for the entire execution, generating a kernel interrupt on every symbol hit. In contrast, single-snapshot mode disables the breakpoint after the first hit, eliminating subsequent interrupt overhead.

The implementation in src/trace.ml (lines 992-998) controls this behavior through the single_hit parameter:

let single_hit = not opts.multi_snapshot in
...
Breakpoint.enable bp ~single_hit |> Or_error.ok_exn;

Because single_hit evaluates to false in multi-snapshot mode, the CPU must context-switch into the kernel on every trigger invocation to handle the breakpoint exception, adding to the observed latency penalty.

Backend Compatibility Limitations

Multi-snapshot mode imposes restrictions on the perf backend command configuration. Specifically, it is incompatible with the exit-based snapshot trigger (snapshot-on-exit).

The src/perf_tool_backend.ml file (lines 167-174) explicitly validates this constraint by checking that users do not combine these incompatible options. When multi_snapshot is enabled, the exit-based snapshot path is disabled because it cannot guarantee correct ordering with continuous collection.

Consequently, multi-snapshot mode only works with function-call triggers (e.g., SIGUSR2 signals or symbol breakpoints) and cannot be combined with process-exit capture points.

Measuring the Performance Impact

To quantify the overhead in your specific environment, compare execution times with and without the flag. The following OCaml benchmark demonstrates the cumulative effect of the ~8 µs penalty over 10,000 iterations:

let run () =
  let start = Unix.gettimeofday () in
  for i = 1 to 10_000 do
    My_module.my_hot_function ()
  done;
  Printf.printf "Elapsed: %f s\n" (Unix.gettimeofday () -. start)

Run the program twice—once with -multi-snapshot enabled and once without—to observe the latency difference directly attributable to trigger overhead.

Practical Usage Examples

Command-Line Configuration

To capture every invocation of my_fun with a limited ring-buffer size per thread:

magic-trace \
  -program ./my_app \
  -trigger my_fun \
  -multi-snapshot \
  -snapshot-size 4M

OCaml API Integration

When using magic-trace as a library, configure the trigger and enable multi-snapshot mode programmatically:

open Magic_trace

let () =
  (* Register a user-defined trigger *)
  let trigger = Magic_trace.Trigger.of_function "my_fun" in
  Magic_trace.set_trigger trigger;

  (* Enable multi-snapshot mode *)
  Magic_trace.enable_multi_snapshot ();

  (* Snapshots now occur automatically on every trigger hit *)
  ignore (Sys.getenv_opt "MAGIC_TRACE_RUN")

When to Use Multi-Snapshot Mode

Despite the performance costs, multi-snapshot mode is invaluable for debugging tight loops or high-frequency events where you need temporal visibility into how behavior evolves over time. Use it when:

  • The trigger fires infrequently enough that 8 µs per hit is negligible
  • You need to correlate state changes across multiple trigger events
  • You can tolerate large trace files and have sufficient disk I/O bandwidth

Avoid multi-snapshot mode for high-throughput paths where microsecond-level latency matters, or when storage and memory constraints prevent handling multi-gigabyte trace files.

Summary

  • Latency cost: Each trigger hit incurs approximately 8 µs of overhead due to buffer flushing and breakpoint re-arming, as documented in src/trace.ml.
  • File size impact: Trace files grow linearly with the number of snapshots, potentially consuming significant memory and disk space via src/trace_writer.ml.
  • Interrupt overhead: The breakpoint remains active, causing kernel interrupts on every trigger invocation (controlled in src/trace.ml).
  • Backend restrictions: Multi-snapshot mode is incompatible with snapshot-on-exit and requires function-call triggers (enforced in src/perf_tool_backend.ml).
  • Use case: Best suited for low-frequency triggers or deep debugging scenarios where temporal continuity outweighs latency costs.

Frequently Asked Questions

How much overhead does multi-snapshot mode add per trigger hit?

Magic-trace's multi-snapshot mode adds approximately 8 microseconds of latency each time the trigger symbol is hit. This cost is documented in the flag definition within src/trace.ml (lines 538-447) and stems from stopping the perf ring-buffer, flushing events, writing metadata, and re-arming the breakpoint.

Why does multi-snapshot mode cause trace files to grow linearly?

Each snapshot captures a complete copy of the current perf buffer state plus metadata. In multi-snapshot mode, these buffers are written sequentially to disk without overwriting previous data, causing the output file size to increase linearly with the number of trigger hits. This behavior is handled by the trace writer logic in src/trace_writer.ml.

Can I use multi-snapshot mode with snapshot-on-exit?

No. According to the backend implementation in src/perf_tool_backend.ml (lines 167-174), multi-snapshot mode is explicitly disabled when snapshot-on-exit is requested because the exit-based path cannot guarantee correct snapshot ordering with continuous collection. Multi-snapshot only works with function-call triggers.

When should I choose single-snapshot over multi-snapshot mode?

Use single-snapshot mode (the default) when tracing high-frequency hot paths where 8 µs of overhead per call would materially impact application performance, or when you need to minimize trace file size. Use multi-snapshot mode only when you need continuous visibility into a specific code region and can tolerate the latency and storage costs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →