# How Zephyr Detects and Normalizes OOM Conditions in Stack Traces for Marin

> Learn how Zephyr detects and normalizes OOM conditions in stack traces for Marin. Scrubbing memory addresses and file paths with regex ensures accurate error reporting.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-28

---

**Zephyr detects OOM conditions by monitoring subprocess return codes for SIGKILL (-9) and normalizes stack traces by scrubbing memory addresses and file paths with regex before propagating MemoryError to the Marin coordinator.**

The **Marin** data pipeline framework (from `marin-community/marin`) relies on **Zephyr** to execute computation shards in isolated Python subprocesses. To handle kernel-level out-of-memory kills gracefully, Zephyr implements a detection and normalization pipeline that converts raw OS signals into deterministic, deduplicable error reports. This system ensures that OOM crashes in worker processes are captured via **faulthandler**, identified by their exit status, sanitized of volatile memory addresses, and classified as infrastructure failures by the coordinator.

## Subprocess Isolation and Crash Capture Architecture

### Isolated Execution with SubprocessRunner

In [`lib/zephyr/src/zephyr/runners.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/runners.py), the `SubprocessRunner` class launches each shard in a separate Python process using the `-u` flag for unbuffered I/O. This ensures that crash output reaches the parent process immediately without buffering delays that could corrupt stack trace data.

### faulthandler Integration for Signal Capture

Inside the child process, located in [`lib/zephyr/src/zephyr/shard_subprocess.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/shard_subprocess.py), the `configure_logging` function activates Python’s built-in `faulthandler` module at startup. This handler dumps the full Python stack trace to **stderr** upon receiving fatal signals including SIGSEGV, SIGABRT, SIGBUS, and critically, upon **SIGKILL** induced by the Linux OOM killer. The parent process captures this output stream for downstream processing.

## Detecting OOM Kills via OS Return Codes

### SIGKILL Detection Logic

After the child process exits, the parent inspects the operating system return code in the `_child_returncode` method. When the Linux OOM killer terminates a process, it sends **SIGKILL**, which manifests as a negative return code of `-signal.SIGKILL` (or `-9`).

### Raising MemoryError for OOM Conditions

When Zephyr detects `returncode == -signal.SIGKILL`, it raises a `MemoryError` with a descriptive message indicating the subprocess was likely OOM-killed. This converts the low-level OS signal into a catchable Python exception that the coordinator can classify distinctly from user-code exceptions.

```python

# Excerpt from lib/zephyr/src/zephyr/runners.py

if returncode == -signal.SIGKILL:          # ‑9 → OOM‑killer

    raise MemoryError(
        f"Subprocess for shard {task.shard_idx} was killed by SIGKILL "
        f"(returncode {returncode}); most likely OOM‑killed by the kernel."
    )

```

## Normalizing Stack Traces for Deterministic Reporting

### Regex-Based Address and Path Scrubbing

Raw **faulthandler** output contains volatile hexadecimal memory addresses and absolute file paths that differ across runs and machines. In [`lib/zephyr/src/zephyr/shard_subprocess.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/shard_subprocess.py), the `normalize_traceback` function uses compiled regular expressions to replace these variables with stable tokens, yielding a reproducible error signature.

```python

# Excerpt from lib/zephyr/src/zephyr/shard_subprocess.py

def normalize_traceback(tb: str) -> str:
    # Replace raw addresses and absolute paths with placeholders

    tb = re.sub(r"0x[0-9a-f]+", "<addr>", tb)
    tb = re.sub(r"/[^\\s]+", "<path>", tb)
    return tb

```

### Serialization Before Coordinator Propagation

The normalized traceback string is pickle-serialized and transmitted to the Marin coordinator. This determinism allows the system to deduplicate identical OOM crashes across different workers and execution attempts, preventing log pollution from memory layout variations.

## Coordinator-Level Infrastructure Failure Handling

### Classifying OOM as Infrastructure Failures

In [`lib/zephyr/src/zephyr/coordinator.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/coordinator.py), the `_record_shard_failure` method receives the propagated `MemoryError` and classifies it as `ShardFailureKind.INFRA` (infrastructure failure) rather than a user-code logic error. This distinction prevents OOM conditions from consuming user-configured retry budgets meant for transient data errors.

### Abort Thresholds for Deterministic Crashers

The coordinator maintains a per-shard counter of infrastructure failures. Once the count exceeds `MAX_SHARD_INFRA_FAILURES`, the coordinator treats the shard as a deterministic crasher—likely due to native SIGSEGV or OOM conditions—and aborts the entire pipeline to prevent infinite retry loops that waste compute resources.

```python

# Excerpt from lib/zephyr/src/zephyr/coordinator.py

if kind is ShardFailureKind.INFRA:
    run.task_infra_attempts[shard_idx] += 1
    if infra_attempts >= self._max_shard_infra_failures:
        logger.error(
            "[%s] Shard %d has been in flight during %d infra failures (max %d). "
            "treating as a deterministic crasher (likely native SIGSEGV / OOM in shard "
            "code) and aborting pipeline. Last failure on worker %s.",
            run.execution_id,
            shard_idx,
            infra_attempts,
            self._max_shard_infra_failures,
            worker_id,
        )
        run.fatal_error = (
            f"Shard {shard_idx} crashed its worker {infra_attempts} times "
            f"(max {self._max_shard_infra_failures} infra failures while in flight). "
            f"Last failure on worker {worker_id}."
        )

```

## Summary

- **SubprocessRunner** executes shards in isolated processes via [`lib/zephyr/src/zephyr/runners.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/runners.py), capturing stderr for crash analysis.
- **OOM detection** occurs by checking for `returncode == -signal.SIGKILL` (-9) in `_child_returncode`, raising `MemoryError` when found.
- **Stack trace normalization** uses regex in `normalize_traceback` (from [`shard_subprocess.py`](https://github.com/marin-community/marin/blob/main/shard_subprocess.py)) to replace memory addresses with `<addr>` and paths with `<path>`.
- **Coordinator logic** in [`lib/zephyr/src/zephyr/coordinator.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/coordinator.py) classifies these errors as `ShardFailureKind.INFRA` and aborts pipelines that exceed `MAX_SHARD_INFRA_FAILURES`.

## Frequently Asked Questions

### How does Zephyr distinguish OOM kills from regular Python exceptions?

Zephyr examines the OS return code in the `_child_returncode` method of `SubprocessRunner`. A return code equal to `-signal.SIGKILL` (-9) specifically indicates the Linux OOM killer terminated the process, prompting Zephyr to raise a `MemoryError` rather than a standard exception.

### What normalization techniques does Zephyr apply to OOM stack traces?

The `normalize_traceback` function in [`lib/zephyr/src/zephyr/shard_subprocess.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/shard_subprocess.py) applies regular expressions to replace hexadecimal memory addresses (matching `0x[0-9a-f]+`) with `<addr>` and absolute file paths (matching `/[^\\s]+`) with `<path>`, creating deterministic signatures suitable for deduplication.

### How does the Marin coordinator handle repeated OOM conditions?

The coordinator classifies OOM-induced `MemoryError` instances as `ShardFailureKind.INFRA`. After the infrastructure failure count for a shard exceeds `MAX_SHARD_INFRA_FAILURES`, the coordinator aborts the entire pipeline, treating the shard as a deterministic crasher to prevent infinite retries.

### Where is the faulthandler configured to capture OOM stack traces?

Inside [`lib/zephyr/src/zephyr/shard_subprocess.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/shard_subprocess.py), the `configure_logging` function enables Python’s `faulthandler` at process startup, ensuring stack traces are written to stderr on SIGKILL and other fatal signals before the process terminates.