How LoopX Session Recovery Works After a Host Failure: A Deep Dive

LoopX treats host failures as recoverable turns by recording a resume_session recovery entry in the turn journal, validating it through schema and session binding checks, and delegating to ChatRuntime.resume_session() to reload persisted state and continue processing without re-executing the failed host step.

LoopX's session recovery mechanism ensures that when a host crashes during code execution, no user progress is lost. Unlike simpler retry systems that re-run operations, LoopX preserves session state and resumes exactly where the failure occurred. This article examines the eight-step recovery flow implemented in the huangruiteng/loopx codebase, with specific source file references and runnable examples.


How Host Failures Are Detected and Recorded

When a host crashes during the host_execute phase, LoopX captures this failure through a structured receipt and host-recovery entry.

The detection happens in loopx/main/loopx/control_plane/turn_driver/session_recovery.py (lines 73–98). The turn driver writes a receipt containing failed_phase: "host_execute" and stores a host_recovery record with two critical fields: schema_version: "loopx_turn_host_recovery_v0" and kind: "resume_session".


# Example: Host-recovery record structure after a crash

receipt = {
    "failed_phase": "host_execute",
    "error": "Segmentation fault",
}
journal = {
    "host_recovery": {
        "schema_version": "loopx_turn_host_recovery_v0",
        "kind": "resume_session",
    },
    "receipt": receipt,
}

This journaling approach distinguishes LoopX's session recovery from blind retry logic—the recovery is intentionally typed and versioned for forward compatibility.


Validating the Recovery Request

The reconcile_failed_turn_retry_request() function orchestrates validation, delegating to assess_failed_turn_retry_request() for schema and kind checking in session_recovery.py (lines 82–99).

Validation Checks Performed

  • Schema version: Must match loopx_turn_host_recovery_v0
  • Kind: Must be resume_session
  • Failed phase: Must be host_execute

Any validation failure raises ValueError, aborting recovery and preventing corrupted or mismatched recovery records from propagating.

from loopx.control_plane.turn_driver.session_recovery import (
    reconcile_failed_turn_retry_request,
)

def retry_failed_turn(request, journal, binding_resolver):
    """Returns a resolved request for session recovery, or raises on invalid input."""
    return reconcile_failed_turn_retry_request(
        request,
        journal,
        session_binding_resolver=binding_resolver,
    )

Resolving and Verifying Session Bindings

After validation, LoopX must ensure the recovered session matches the active session. This prevents session hijacking or accidental cross-session recovery.

Step 4: Extract the Original Binding

A session-binding resolver (injected by the turn driver) extracts the original session binding from the turn envelope, describing which LoopX session the failed turn belongs to.

Step 5: Match Against Live Session

The reconcile_failed_turn_session_request() function (imported from driver.py and called in session_recovery.py lines 113–119) compares the resolved binding against the live session in the state store. Mismatched IDs raise FailedTurnSessionRecoveryError, aborting recovery.

This binding verification is critical for multi-tenant or long-running deployments where session ID confusion could cause data corruption.


Constructing the Resolved Recovery Request

Once binding verification passes, assess_failed_turn_retry_request() returns a resolved request containing:

  • The original payload
  • A flag indicating the host should be skipped on retry

This flag prevents re-execution of the failed host step, avoiding repeated crashes from deterministic failures.


# The resolved request structure (conceptual)

resolved_request = {
    "session": {
        "action": "resume",  # Changed from "start_new"

        "id": "sess_abc123"
    },
    "turn_envelope": {...},
    "_recovery": {
        "skip_host": True,  # Critical: don't re-run the host step

    }
}

Runtime Session Resumption

The final phase delegates to ChatRuntime.resume_session() in loopx/main/loopx/chat_runtime.py. This method performs the actual state restoration.


# From loopx/chat_runtime.py

def resume_session(self, *, session_id: str, work_dir: Path, objective: str) -> dict:
    """Load the persisted session and continue processing its queue."""
    session = self.store.load_session(session_id)
    if not session:
        raise KeyError("session not found")
    
    # Re-create the host adapter if needed

    adapter = self._ensure_adapter(session, work_dir=work_dir, objective=objective)
    
    # Drain queued turns—host step is now skipped per recovery flags

    self._drain_session_queue(
        session_id=session_id,
        work_dir=work_dir,
        objective=objective
    )
    return session

The runtime handles three responsibilities:

  1. State hydration: Loads session from self.store.load_session()
  2. Adapter reconstruction: Re-establishes the host environment via _ensure_adapter()
  3. Queue processing: Resumes turn execution through _drain_session_queue()

Key Source Files and Their Roles

File Purpose
loopx/main/loopx/control_plane/turn_driver/session_recovery.py Implements host-recovery validation, schema checking, and retry request construction
loopx/main/loopx/control_plane/turn_driver/codex_cli.py Sets recovery_kind="resume_session" when observing host failures
loopx/main/loopx/chat_runtime.py Contains resume_session() and queue-draining logic for state restoration
loopx/main/tests/test_loopx_turn_executor.py Tests the resume_session recovery path (lines ~940–960)
loopx/main/tests/test_loopx_turn_driver.py Validates session binding mismatch handling

Testing the Recovery Path

The test suite in test_loopx_turn_executor.py verifies end-to-end recovery behavior.

def test_run_once_resumes_session_observed_by_recoverable_failed_turn():
    # Arrange: Turn fails in host phase with resume_session recovery

    request = {"session": {"action": "start_new"}, "turn_envelope": {...}}
    journal = {
        "host_recovery": {
            "schema_version": "loopx_turn_host_recovery_v0",
            "kind": "resume_session"
        },
        "receipt": {"failed_phase": "host_execute"},
    }
    
    # Act: Driver processes the failed turn

    recovered = reconcile_failed_turn_retry_request(
        request, 
        journal, 
        session_binding_resolver=my_resolver
    )
    
    # Assert: Recovery switches action to "resume"

    assert recovered["session"]["action"] == "resume"

Additional tests in test_loopx_turn_driver.py confirm that missing or mismatched session bindings properly abort recovery with FailedTurnSessionRecoveryError.


Summary

  • Host failures in LoopX are captured as typed resume_session recovery entries in the turn journal
  • Validation layers (schema, kind, failed_phase) prevent corrupted recovery attempts
  • Session binding resolution ensures recovered sessions match the active session, preventing cross-session contamination
  • Runtime resumption via ChatRuntime.resume_session() reloads persisted state without re-executing the failed host step
  • Skip-host semantics guarantee progress for deterministic failures while preserving the execution context

Frequently Asked Questions

What happens if the session binding doesn't match during recovery?

LoopX raises FailedTurnSessionRecoveryError and aborts recovery. This safety mechanism in session_recovery.py (lines 113–119) prevents accidental session hijacking or data corruption from stale or tampered recovery records.

Does LoopX re-execute the host code after a crash?

No. The resolved recovery request includes a skip_host flag that prevents re-execution. The runtime loads persisted session state through resume_session() and continues from the next queued turn, preserving all prior progress.

What schema versions does LoopX support for host recovery?

As of the analyzed codebase, LoopX uses loopx_turn_host_recovery_v0 with kind: "resume_session". The assess_failed_turn_retry_request() function validates both fields, and unsupported schemas trigger ValueError with no retry attempted.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →