How Equilibrium Engine Implements Crash Protection on Windows for SEGFAULT and SIGABRT

Equilibrium Engine prevents host process termination on Windows by intercepting both access violations (SEGFAULT) and abort signals (SIGABRT) through the embedded cr library, using Structured Exception Handling (SEH) and setjmp/longjmp to roll back to the last stable version of a hot-reloaded plugin.

The engine's live-reload system allows developers to iterate on DLLs without restarting the application. According to the source code in clibequilibrium/equilibriumengine, this crash protection mechanism ensures that bugs in recompiled code—such as null pointer dereferences or failed assertions—are isolated to the plugin boundary, allowing the host process in launcher/launcher.cc to continue running uninterrupted.

Windows Crash Handling Architecture

On Windows, the crash protection system leverages two distinct recovery strategies depending on the failure type. The implementation resides entirely in 3rdparty/cr/cr.h and is activated through cr_plat_init() at the start of each frame in the host application.

  • SEGFAULT (EXCEPTION_ACCESS_VIOLATION): Caught by the SEH __except filter, translating the Windows exception code into CR_SEGFAULT and triggering a rollback to ctx.last_working_version.
  • SIGABRT: Caught by a custom signal() handler that uses setjmp/longjmp to perform a non-local jump back to the host's control flow without terminating the process.

SIGABRT Recovery via Signal Handlers

For explicit program aborts triggered by abort() or failed assertions, the library installs a custom signal handler that transforms the fatal signal into a recoverable jump. In 3rdparty/cr/cr.h, the cr_signal_handler function breaks into the debugger if attached, then performs a longjmp to a previously established recovery point:

static void cr_signal_handler(int sig) {
    __debugbreak();                // Breaks into debugger (if attached)
    longjmp(env, 1);               // Jump back to the setjmp point
}

The handler is registered during platform initialization via signal(SIGABRT, cr_signal_handler) inside cr_plat_init(). When a plugin calls abort(), control returns to the setjmp location in cr_plugin_main rather than crashing the host.

SEGFAULT Recovery via Structured Exception Handling (SEH)

For memory access violations, the engine uses Windows Structured Exception Handling. The cr_seh_filter function translates Windows exception codes into the library's internal cr_failure enum and determines whether the host can safely recover:

static int cr_seh_filter(cr_plugin &ctx, unsigned long seh) {
    if (ctx.version == 1) {
        // First version cannot rollback, continue search for debugger
        return EXCEPTION_CONTINUE_SEARCH;
    }

    ctx.version = ctx.last_working_version;  // Perform rollback

    switch (seh) {
        case EXCEPTION_ACCESS_VIOLATION: 
            ctx.failure = CR_SEGFAULT; 
            break;
        case EXCEPTION_ILLEGAL_INSTRUCTION: 
            ctx.failure = CR_ILLEGAL; 
            break;
    }
    
    return EXCEPTION_EXECUTE_HANDLER;
}

This filter runs inside the __except clause of cr_plugin_main. When EXCEPTION_ACCESS_VIOLATION is detected, the filter assigns CR_SEGFAULT to ctx.failure and triggers a rollback, preventing the crash from propagating to the operating system.

The Plugin Execution Wrapper

The cr_plugin_main function in 3rdparty/cr/cr.h orchestrates both protection mechanisms. It establishes a recovery point with setjmp before entering a __try block that executes the plugin's entry point:

static int cr_plugin_main(cr_plugin &ctx, cr_op operation) {
    auto p = (cr_internal *)ctx.p;

    if (int sig = setjmp(env)) {               // Jump here on SIGABRT
        ctx.version = ctx.last_working_version;
        ctx.failure = cr_signal_to_failure(sig);
        return -1;                              // Signal-induced failure
    } else {
        __try {
            if (p->main) {
                return p->main(&ctx, operation);
            }
        } __except (cr_seh_filter(ctx, GetExceptionCode())) {
            return -1;                          // SEH-induced failure
        }
    }
    return -1;
}

Key implementation detail: The combination of setjmp/longjmp for signals and __try/__except for SEH guarantees that any fatal Windows error originating from the plugin is caught. The function returns -1 to indicate failure, while the host process remains stable.

Host Integration and Fallback Flow

The host application in launcher/launcher.cc drives the protection system by calling cr_plugin_update within its main loop. This function internally invokes cr_plugin_main and checks the return value to detect crashes:

while (engine.running) {
    if (cr_plugin_update(editor) == 0 && cr_plugin_update(application) == 0) {
        cr_plugin_update(engine_simulation);
    }
    cr_plat_init();               // (re)install SIGABRT handler each frame
}

When a plugin crashes—whether through a SEGFAULT or SIGABRT—cr_plugin_update returns -1. The host simply skips that plugin's update for the current frame and continues execution. The library automatically maintains ctx.last_working_version, ensuring that subsequent reload attempts will use the last known good DLL until the developer provides a fixed build.

Practical Example: Simulating a Crash

To verify the protection mechanism, a developer can deliberately trigger a crash in a hot-reloadable module. Consider a simplified plugin:

CR_EXPORT int cr_main(struct cr_plugin *ctx, enum cr_op operation) {
    if (operation == CR_STEP) {
        if (should_test_abort) {
            abort();          // Triggers SIGABRT → cr_signal_handler → longjmp
        }
        // Or: int* p = NULL; *p = 42;  // Would trigger SEGFAULT → SEH filter
    }
    return 0;
}

When running the host from launcher/launcher.cc:

  1. The host loads the plugin DLL and executes cr_main.
  2. Upon reaching abort(), the registered cr_signal_handler intercepts the signal.
  3. The handler invokes longjmp, returning control to cr_plugin_main with a non-zero value.
  4. The engine sets ctx.failure to CR_ABORT, rolls back ctx.version, and returns -1.
  5. The host loop continues; the process survives, and a fixed DLL can be hot-reloaded immediately.

Summary

  • Dual mechanism: Windows crash protection in Equilibrium Engine combines SIGABRT handling via signal() and longjmp with SEGFAULT handling via SEH __try/__except blocks, both implemented in 3rdparty/cr/cr.h.
  • Automatic rollback: When cr_seh_filter catches EXCEPTION_ACCESS_VIOLATION or cr_signal_handler catches an abort signal, the engine reverts ctx.version to ctx.last_working_version, isolating the crash to the plugin.
  • Host resilience: The main loop in launcher/launcher.cc invokes cr_plat_init() each frame and checks cr_plugin_update return values, ensuring the host process never terminates due to plugin errors.
  • Developer workflow: Developers can safely test edge cases and assertions knowing that crashes trigger an immediate fallback to the last stable DLL without restarting the engine.

Frequently Asked Questions

Does the crash protection work with any Windows compiler?

The current implementation in 3rdparty/cr/cr.h specifically checks for __clang__ when registering signal handlers and establishing setjmp recovery points. While the SEH filter works with standard Windows exception handling, the SIGABRT protection via signal() and longjmp is wrapped in Clang-specific preprocessor guards, indicating this configuration is the primary target for Windows builds in this repository.

What happens if the very first version of a plugin crashes?

In cr_seh_filter, the code explicitly checks if (ctx.version == 1) and returns EXCEPTION_CONTINUE_SEARCH for initial loads. This means if the first version of a DLL crashes upon loading, the exception propagates to the debugger or operating system rather than attempting a rollback to a non-existent previous version. Rollback protection activates only after at least one successful version has been established.

How does the host detect that a plugin has crashed?

The host detects crashes through the return value of cr_plugin_update, which internally calls cr_plugin_main. If either the SEH filter catches an access violation or the signal handler catches an abort, the function returns -1. The host loop in launcher/launcher.cc checks this value and skips further processing of that plugin for the current frame while keeping the application running.

Can the protection handle stack overflow or other critical errors?

The documented implementation specifically handles EXCEPTION_ACCESS_VIOLATION for SEGFAULTS and SIGABRT for assertion failures. While the SEH filter in cr_seh_filter includes cases for EXCEPTION_ILLEGAL_INSTRUCTION and other codes, severe conditions like stack overflow (EXCEPTION_STACK_OVERFLOW) would likely still terminate the process, as the recovery mechanism requires sufficient stack space to execute the filter and rollback logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →