# How LoopX Session Recovery Works After a Host Failure: A Deep Dive

> Discover how LoopX handles host failures with its session recovery mechanism. Learn how LoopX recovers sessions, validates data, and resumes processing without re-execution.

- Repository: [huangruiteng/loopx](https://github.com/huangruiteng/loopx)
- Tags: deep-dive
- Published: 2026-09-04

---

**LoopX treats host failures as recoverable turns by recording a `resume_session` recovery entry in the turn journal, validating it through schema and session binding checks, and delegating to `ChatRuntime.resume_session()` to reload persisted state and continue processing without re-executing the failed host step.**

LoopX's **session recovery** mechanism ensures that when a host crashes during code execution, no user progress is lost. Unlike simpler retry systems that re-run operations, LoopX preserves session state and resumes exactly where the failure occurred. This article examines the eight-step recovery flow implemented in the `huangruiteng/loopx` codebase, with specific source file references and runnable examples.

---

## How Host Failures Are Detected and Recorded

When a host crashes during the `host_execute` phase, LoopX captures this failure through a structured **receipt** and **host-recovery entry**.

The detection happens in [`loopx/main/loopx/control_plane/turn_driver/session_recovery.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/main/loopx/control_plane/turn_driver/session_recovery.py) (lines 73–98). The turn driver writes a receipt containing `failed_phase: "host_execute"` and stores a `host_recovery` record with two critical fields: `schema_version: "loopx_turn_host_recovery_v0"` and `kind: "resume_session"`.

```python

# Example: Host-recovery record structure after a crash

receipt = {
    "failed_phase": "host_execute",
    "error": "Segmentation fault",
}
journal = {
    "host_recovery": {
        "schema_version": "loopx_turn_host_recovery_v0",
        "kind": "resume_session",
    },
    "receipt": receipt,
}

```

This journaling approach distinguishes LoopX's **session recovery** from blind retry logic—the recovery is intentionally typed and versioned for forward compatibility.

---

## Validating the Recovery Request

The `reconcile_failed_turn_retry_request()` function orchestrates validation, delegating to `assess_failed_turn_retry_request()` for schema and kind checking in [`session_recovery.py`](https://github.com/huangruiteng/loopx/blob/main/session_recovery.py) (lines 82–99).

### Validation Checks Performed

- **Schema version**: Must match `loopx_turn_host_recovery_v0`
- **Kind**: Must be `resume_session`
- **Failed phase**: Must be `host_execute`

Any validation failure raises `ValueError`, aborting recovery and preventing corrupted or mismatched recovery records from propagating.

```python
from loopx.control_plane.turn_driver.session_recovery import (
    reconcile_failed_turn_retry_request,
)

def retry_failed_turn(request, journal, binding_resolver):
    """Returns a resolved request for session recovery, or raises on invalid input."""
    return reconcile_failed_turn_retry_request(
        request,
        journal,
        session_binding_resolver=binding_resolver,
    )

```

---

## Resolving and Verifying Session Bindings

After validation, LoopX must ensure the recovered session matches the active session. This prevents **session hijacking** or accidental cross-session recovery.

### Step 4: Extract the Original Binding

A **session-binding resolver** (injected by the turn driver) extracts the original session binding from the turn envelope, describing which LoopX session the failed turn belongs to.

### Step 5: Match Against Live Session

The `reconcile_failed_turn_session_request()` function (imported from [`driver.py`](https://github.com/huangruiteng/loopx/blob/main/driver.py) and called in [`session_recovery.py`](https://github.com/huangruiteng/loopx/blob/main/session_recovery.py) lines 113–119) compares the resolved binding against the live session in the state store. Mismatched IDs raise `FailedTurnSessionRecoveryError`, aborting recovery.

This binding verification is critical for multi-tenant or long-running deployments where session ID confusion could cause data corruption.

---

## Constructing the Resolved Recovery Request

Once binding verification passes, `assess_failed_turn_retry_request()` returns a **resolved request** containing:

- The original payload
- A flag indicating the host should be **skipped** on retry

This flag prevents re-execution of the failed host step, avoiding repeated crashes from deterministic failures.

```python

# The resolved request structure (conceptual)

resolved_request = {
    "session": {
        "action": "resume",  # Changed from "start_new"

        "id": "sess_abc123"
    },
    "turn_envelope": {...},
    "_recovery": {
        "skip_host": True,  # Critical: don't re-run the host step

    }
}

```

---

## Runtime Session Resumption

The final phase delegates to `ChatRuntime.resume_session()` in [`loopx/main/loopx/chat_runtime.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/main/loopx/chat_runtime.py). This method performs the actual state restoration.

```python

# From loopx/chat_runtime.py

def resume_session(self, *, session_id: str, work_dir: Path, objective: str) -> dict:
    """Load the persisted session and continue processing its queue."""
    session = self.store.load_session(session_id)
    if not session:
        raise KeyError("session not found")
    
    # Re-create the host adapter if needed

    adapter = self._ensure_adapter(session, work_dir=work_dir, objective=objective)
    
    # Drain queued turns—host step is now skipped per recovery flags

    self._drain_session_queue(
        session_id=session_id,
        work_dir=work_dir,
        objective=objective
    )
    return session

```

The runtime handles three responsibilities:

1. **State hydration**: Loads session from `self.store.load_session()`
2. **Adapter reconstruction**: Re-establishes the host environment via `_ensure_adapter()`
3. **Queue processing**: Resumes turn execution through `_drain_session_queue()`

---

## Key Source Files and Their Roles

| File | Purpose |
|------|---------|
| [`loopx/main/loopx/control_plane/turn_driver/session_recovery.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/main/loopx/control_plane/turn_driver/session_recovery.py) | Implements host-recovery validation, schema checking, and retry request construction |
| [`loopx/main/loopx/control_plane/turn_driver/codex_cli.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/main/loopx/control_plane/turn_driver/codex_cli.py) | Sets `recovery_kind="resume_session"` when observing host failures |
| [`loopx/main/loopx/chat_runtime.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/main/loopx/chat_runtime.py) | Contains `resume_session()` and queue-draining logic for state restoration |
| [`loopx/main/tests/test_loopx_turn_executor.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/main/tests/test_loopx_turn_executor.py) | Tests the resume_session recovery path (lines ~940–960) |
| [`loopx/main/tests/test_loopx_turn_driver.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/main/tests/test_loopx_turn_driver.py) | Validates session binding mismatch handling |

---

## Testing the Recovery Path

The test suite in [`test_loopx_turn_executor.py`](https://github.com/huangruiteng/loopx/blob/main/test_loopx_turn_executor.py) verifies end-to-end recovery behavior.

```python
def test_run_once_resumes_session_observed_by_recoverable_failed_turn():
    # Arrange: Turn fails in host phase with resume_session recovery

    request = {"session": {"action": "start_new"}, "turn_envelope": {...}}
    journal = {
        "host_recovery": {
            "schema_version": "loopx_turn_host_recovery_v0",
            "kind": "resume_session"
        },
        "receipt": {"failed_phase": "host_execute"},
    }
    
    # Act: Driver processes the failed turn

    recovered = reconcile_failed_turn_retry_request(
        request, 
        journal, 
        session_binding_resolver=my_resolver
    )
    
    # Assert: Recovery switches action to "resume"

    assert recovered["session"]["action"] == "resume"

```

Additional tests in [`test_loopx_turn_driver.py`](https://github.com/huangruiteng/loopx/blob/main/test_loopx_turn_driver.py) confirm that missing or mismatched session bindings properly abort recovery with `FailedTurnSessionRecoveryError`.

---

## Summary

- **Host failures** in LoopX are captured as typed `resume_session` recovery entries in the turn journal
- **Validation layers** (schema, kind, failed_phase) prevent corrupted recovery attempts
- **Session binding resolution** ensures recovered sessions match the active session, preventing cross-session contamination
- **Runtime resumption** via `ChatRuntime.resume_session()` reloads persisted state without re-executing the failed host step
- **Skip-host semantics** guarantee progress for deterministic failures while preserving the execution context

---

## Frequently Asked Questions

### What happens if the session binding doesn't match during recovery?

LoopX raises `FailedTurnSessionRecoveryError` and aborts recovery. This safety mechanism in [`session_recovery.py`](https://github.com/huangruiteng/loopx/blob/main/session_recovery.py) (lines 113–119) prevents accidental session hijacking or data corruption from stale or tampered recovery records.

### Does LoopX re-execute the host code after a crash?

No. The resolved recovery request includes a `skip_host` flag that prevents re-execution. The runtime loads persisted session state through `resume_session()` and continues from the next queued turn, preserving all prior progress.

### What schema versions does LoopX support for host recovery?

As of the analyzed codebase, LoopX uses `loopx_turn_host_recovery_v0` with `kind: "resume_session"`. The `assess_failed_turn_retry_request()` function validates both fields, and unsupported schemas trigger `ValueError` with no retry attempted.