# Magic-Trace Multi-Thread Recording and Snapshot Buffer Allocation Considerations

> Understand magic-trace multi-thread recording and snapshot buffer allocation. Learn how the buffer divides equally for power-of-two page alignment and its impact on look-back windows.

- Repository: [Jane Street/magic-trace](https://github.com/janestreet/magic-trace)
- Tags: deep-dive
- Published: 2026-05-24

---

**When multi-thread recording is enabled in magic-trace, the finite snapshot buffer is divided equally among all threads, reducing per-thread look-back windows while enforcing power-of-two page alignment constraints.**

Magic-trace, the time-travel debugging tool developed by Jane Street, provides powerful Intel PT-based tracing capabilities for OCaml and other applications. Understanding how the tool allocates snapshot buffers when recording multiple threads is critical for capturing complete execution traces without dropping historical data. The architecture imposes specific constraints on buffer sizing and memory layout that directly impact trace depth and system resource usage.

## How Multi-Thread Recording Divides the Snapshot Buffer

When the `-multi-thread` flag is enabled, the tracer must share the finite snapshot buffer across every thread it records. According to the implementation in [`src/perf_tool_backend.ml`](https://github.com/janestreet/magic-trace/blob/main/src/perf_tool_backend.ml) (lines 18-27), the backend reserves a fixed amount of memory for the ring-buffer that holds snapshot data. Enabling multi-thread mode splits this memory evenly across all recorded threads, effectively reducing the per-thread look-back window proportionally to the thread count.

This architectural decision ensures fairness across threads but introduces a direct trade-off: capturing more threads means retaining less history for each individual thread.

### Default Buffer Sizing Based on Permissions

The default snapshot buffer size varies based on the host's `perf_event_paranoid` setting. As defined in [`src/perf_tool_backend.ml`](https://github.com/janestreet/magic-trace/blob/main/src/perf_tool_backend.ml) (lines 10-15), privileged users receive approximately **4 MiB**, while unprivileged users receive approximately **256 KiB**. Users can override these defaults using the `-snapshot-size` command-line option, which accepts a string representation of bytes (e.g., `"8M"` for 8 megabytes).

## Snapshot Size Constraints and Power-of-Two Alignment

Magic-trace enforces strict memory alignment requirements to simplify address-space calculations and enable efficient `mmap` operations. The `Pow2_pages` module in [`src/pow2_pages.ml`](https://github.com/janestreet/magic-trace/blob/main/src/pow2_pages.ml) handles all snapshot size validation and rounding, ensuring that buffers can be mapped with a single system call.

### Power-of-Two Page Rounding

The **`Pow2_pages.create`** function (lines 15-33) automatically rounds any user-specified size to the nearest power-of-two number of pages. If the requested size requires rounding up or down, the system emits a warning to notify the user of the adjustment. This guarantees that the buffer can be mapped with a single `mmap` call and simplifies circular buffer arithmetic for the ring buffer implementation.

### Maximum Address-Space Limits

The module enforces hard limits on virtual memory consumption to prevent system instability. As implemented in [`src/pow2_pages.ml`](https://github.com/janestreet/magic-trace/blob/main/src/pow2_pages.ml) (lines 9-12), the code caps buffers at the x86-64 virtual address space limit of **2^48 bytes**. If a user requests a snapshot size exceeding this threshold, the system warns about the constraint and clamps the value, preventing address space exhaustion that would otherwise terminate the tracing session.

## Impact on Look-Back Period and Trace Analysis

Because the total buffer size remains fixed regardless of thread count, splitting the buffer across more threads directly shortens the amount of historic data each thread can retain. This trade-off requires careful consideration when analyzing applications with sporadic events or long-running operations.

Users must balance the need for full-thread coverage against the reduced snapshot depth. For analyses requiring deep historical data that exceeds the divided buffer capacity, consider using the `-full-execution` flag, which disables the ring-buffer and records the entire execution trace to disk. While this avoids buffer limitations, it generates massive trace files suitable only for short runs or scenarios with abundant storage.

## Snapshot Trigger Behavior in Multi-Thread Mode

The snapshot trigger mechanism (`-trigger`) operates consistently regardless of thread count, but multi-thread mode affects how triggers interact with the recording lifecycle. According to [`src/trace.ml`](https://github.com/janestreet/magic-trace/blob/main/src/trace.ml) (lines 389-397), when `-multi-thread` is enabled, the tracer may need to take snapshots on every thread that hits the breakpoint.

In **single-snapshot mode**, the breakpoint disables immediately after the first hit to prevent repeated captures. However, in **multi-snapshot mode**, the breakpoint remains enabled across threads, allowing continuous capture throughout the execution lifetime of multiple threads.

## Practical Configuration Examples

Configure multi-thread recording and custom buffer sizes using the following patterns from the `janestreet/magic-trace` source.

### OCaml API Configuration

```ocaml
(* Enable multi-thread recording with 8 MiB snapshot buffer *)
let record_opts =
  Perf_tool_backend.Record_opts.{ 
    multi_thread = true;
    full_execution = false;
    snapshot_size = Some (Pow2_pages.create (Byte_units.of_string "8M"));
    callgraph_mode = None;
  }

```

### Command-Line Interface

```bash

# Record all threads with 8 MiB snapshot buffer

magic-trace run -multi-thread -snapshot-size 8M ./my_program

```

### Runtime Buffer Inspection

```ocaml
(* Calculate effective per-thread buffer size at runtime *)
let per_thread_pages = 
  match record_opts.snapshot_size with
  | Some sz -> Pow2_pages.num_pages sz
  | None   -> failwith "No snapshot size provided"
;;
let bytes_per_page = Int64.of_int 4096 in
let total_bytes = 
  Int64.( * ) (Int64.of_int per_thread_pages) bytes_per_page in
Printf.printf "Each thread allocated: %s\n"
  (Byte_units.to_string_hum (Byte_units.of_bytes_int64_exn total_bytes))

```

## Summary

- Enabling `-multi-thread` divides the total snapshot buffer equally among all recorded threads, reducing individual look-back capacity proportionally to the thread count.
- Buffer sizes must be power-of-two multiples of the system page size, enforced by the `Pow2_pages` module in [`src/pow2_pages.ml`](https://github.com/janestreet/magic-trace/blob/main/src/pow2_pages.ml).
- Hard limits prevent allocation beyond the 2^48 byte x86-64 address space ceiling, with warnings emitted for oversized requests.
- Reduced per-thread buffers may necessitate `-full-execution` mode for deep historical analysis, trading disk space for unlimited trace depth.
- Multi-thread mode modifies breakpoint behavior to allow continuous triggering across threads when multi-snapshot recording is active.

## Frequently Asked Questions

### How does multi-thread mode affect the amount of history captured per thread?

When multi-thread recording is enabled, the fixed snapshot buffer is divided equally among all threads. If you allocate a 16 MiB buffer and trace four threads, each thread retains only approximately 4 MiB of execution history. This linear reduction in per-thread buffer space directly limits how far back in time you can inspect when a trigger fires, making thread count a critical tuning parameter for deep historical analysis.

### Why must snapshot sizes be power-of-two multiples of page size?

The `Pow2_pages` module enforces this constraint to optimize virtual memory mapping and simplify circular buffer address calculations. As implemented in [`src/pow2_pages.ml`](https://github.com/janestreet/magic-trace/blob/main/src/pow2_pages.ml), power-of-two sizing guarantees that the entire buffer can be allocated with a single `mmap` call and enables efficient wrap-around arithmetic for the ring buffer implementation, eliminating the need for expensive modulo operations during trace capture.

### What happens if I request a snapshot size larger than available address space?

The system caps the allocation at 2^48 bytes, the x86-64 architectural limit. The `Pow2_pages` validation logic in [`src/pow2_pages.ml`](https://github.com/janestreet/magic-trace/blob/main/src/pow2_pages.ml) (lines 9-12) detects oversized requests and emits a warning while clamping the value to the maximum safe limit, preventing memory mapping failures that would otherwise terminate the tracing session with an out-of-memory error.

### Can I use the same snapshot size for single-thread and multi-thread recordings?

While you can specify identical `-snapshot-size` values, the effective per-thread history differs significantly. Single-thread mode dedicates the entire buffer to one thread, while multi-thread mode splits it across all recorded threads. For equivalent per-thread history depth in multi-thread mode, multiply your desired per-thread size by the expected thread count before passing it to `-snapshot-size`, keeping in mind the power-of-two rounding requirements and address space limits.