How Zephyr Detects and Normalizes OOM Conditions in Stack Traces for Marin

Zephyr detects OOM conditions by monitoring subprocess return codes for SIGKILL (-9) and normalizes stack traces by scrubbing memory addresses and file paths with regex before propagating MemoryError to the Marin coordinator.

The Marin data pipeline framework (from marin-community/marin) relies on Zephyr to execute computation shards in isolated Python subprocesses. To handle kernel-level out-of-memory kills gracefully, Zephyr implements a detection and normalization pipeline that converts raw OS signals into deterministic, deduplicable error reports. This system ensures that OOM crashes in worker processes are captured via faulthandler, identified by their exit status, sanitized of volatile memory addresses, and classified as infrastructure failures by the coordinator.

Subprocess Isolation and Crash Capture Architecture

Isolated Execution with SubprocessRunner

In lib/zephyr/src/zephyr/runners.py, the SubprocessRunner class launches each shard in a separate Python process using the -u flag for unbuffered I/O. This ensures that crash output reaches the parent process immediately without buffering delays that could corrupt stack trace data.

faulthandler Integration for Signal Capture

Inside the child process, located in lib/zephyr/src/zephyr/shard_subprocess.py, the configure_logging function activates Python’s built-in faulthandler module at startup. This handler dumps the full Python stack trace to stderr upon receiving fatal signals including SIGSEGV, SIGABRT, SIGBUS, and critically, upon SIGKILL induced by the Linux OOM killer. The parent process captures this output stream for downstream processing.

Detecting OOM Kills via OS Return Codes

SIGKILL Detection Logic

After the child process exits, the parent inspects the operating system return code in the _child_returncode method. When the Linux OOM killer terminates a process, it sends SIGKILL, which manifests as a negative return code of -signal.SIGKILL (or -9).

Raising MemoryError for OOM Conditions

When Zephyr detects returncode == -signal.SIGKILL, it raises a MemoryError with a descriptive message indicating the subprocess was likely OOM-killed. This converts the low-level OS signal into a catchable Python exception that the coordinator can classify distinctly from user-code exceptions.


# Excerpt from lib/zephyr/src/zephyr/runners.py

if returncode == -signal.SIGKILL:          # ‑9 → OOM‑killer

    raise MemoryError(
        f"Subprocess for shard {task.shard_idx} was killed by SIGKILL "
        f"(returncode {returncode}); most likely OOM‑killed by the kernel."
    )

Normalizing Stack Traces for Deterministic Reporting

Regex-Based Address and Path Scrubbing

Raw faulthandler output contains volatile hexadecimal memory addresses and absolute file paths that differ across runs and machines. In lib/zephyr/src/zephyr/shard_subprocess.py, the normalize_traceback function uses compiled regular expressions to replace these variables with stable tokens, yielding a reproducible error signature.


# Excerpt from lib/zephyr/src/zephyr/shard_subprocess.py

def normalize_traceback(tb: str) -> str:
    # Replace raw addresses and absolute paths with placeholders

    tb = re.sub(r"0x[0-9a-f]+", "<addr>", tb)
    tb = re.sub(r"/[^\\s]+", "<path>", tb)
    return tb

Serialization Before Coordinator Propagation

The normalized traceback string is pickle-serialized and transmitted to the Marin coordinator. This determinism allows the system to deduplicate identical OOM crashes across different workers and execution attempts, preventing log pollution from memory layout variations.

Coordinator-Level Infrastructure Failure Handling

Classifying OOM as Infrastructure Failures

In lib/zephyr/src/zephyr/coordinator.py, the _record_shard_failure method receives the propagated MemoryError and classifies it as ShardFailureKind.INFRA (infrastructure failure) rather than a user-code logic error. This distinction prevents OOM conditions from consuming user-configured retry budgets meant for transient data errors.

Abort Thresholds for Deterministic Crashers

The coordinator maintains a per-shard counter of infrastructure failures. Once the count exceeds MAX_SHARD_INFRA_FAILURES, the coordinator treats the shard as a deterministic crasher—likely due to native SIGSEGV or OOM conditions—and aborts the entire pipeline to prevent infinite retry loops that waste compute resources.


# Excerpt from lib/zephyr/src/zephyr/coordinator.py

if kind is ShardFailureKind.INFRA:
    run.task_infra_attempts[shard_idx] += 1
    if infra_attempts >= self._max_shard_infra_failures:
        logger.error(
            "[%s] Shard %d has been in flight during %d infra failures (max %d). "
            "treating as a deterministic crasher (likely native SIGSEGV / OOM in shard "
            "code) and aborting pipeline. Last failure on worker %s.",
            run.execution_id,
            shard_idx,
            infra_attempts,
            self._max_shard_infra_failures,
            worker_id,
        )
        run.fatal_error = (
            f"Shard {shard_idx} crashed its worker {infra_attempts} times "
            f"(max {self._max_shard_infra_failures} infra failures while in flight). "
            f"Last failure on worker {worker_id}."
        )

Summary

  • SubprocessRunner executes shards in isolated processes via lib/zephyr/src/zephyr/runners.py, capturing stderr for crash analysis.
  • OOM detection occurs by checking for returncode == -signal.SIGKILL (-9) in _child_returncode, raising MemoryError when found.
  • Stack trace normalization uses regex in normalize_traceback (from shard_subprocess.py) to replace memory addresses with <addr> and paths with <path>.
  • Coordinator logic in lib/zephyr/src/zephyr/coordinator.py classifies these errors as ShardFailureKind.INFRA and aborts pipelines that exceed MAX_SHARD_INFRA_FAILURES.

Frequently Asked Questions

How does Zephyr distinguish OOM kills from regular Python exceptions?

Zephyr examines the OS return code in the _child_returncode method of SubprocessRunner. A return code equal to -signal.SIGKILL (-9) specifically indicates the Linux OOM killer terminated the process, prompting Zephyr to raise a MemoryError rather than a standard exception.

What normalization techniques does Zephyr apply to OOM stack traces?

The normalize_traceback function in lib/zephyr/src/zephyr/shard_subprocess.py applies regular expressions to replace hexadecimal memory addresses (matching 0x[0-9a-f]+) with <addr> and absolute file paths (matching /[^\\s]+) with <path>, creating deterministic signatures suitable for deduplication.

How does the Marin coordinator handle repeated OOM conditions?

The coordinator classifies OOM-induced MemoryError instances as ShardFailureKind.INFRA. After the infrastructure failure count for a shard exceeds MAX_SHARD_INFRA_FAILURES, the coordinator aborts the entire pipeline, treating the shard as a deterministic crasher to prevent infinite retries.

Where is the faulthandler configured to capture OOM stack traces?

Inside lib/zephyr/src/zephyr/shard_subprocess.py, the configure_logging function enables Python’s faulthandler at process startup, ensuring stack traces are written to stderr on SIGKILL and other fatal signals before the process terminates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →