What Is the Purpose of engine-pool.ts in OpenWork? Understanding Blue/Green Engine Management

The engine-pool.ts file implements a blue/green rollover manager that enables zero-downtime configuration updates for the OpenCode engine by orchestrating standby creation, session pinning, and graceful draining of old engine instances.

The engine-pool.ts module serves as the critical infrastructure layer in OpenWork's server architecture, managing the lifecycle of the OpenCode engine processes. Located at apps/server/src/engine-pool.ts, this file ensures that configuration changes—such as model list updates or workspace modifications—can be applied without interrupting active user sessions. By maintaining a sophisticated state machine and request routing system, the engine pool balances operational stability with rapid configuration agility.

Core Responsibilities of engine-pool.ts

The engine pool handles seven distinct operational responsibilities that collectively guarantee high availability. Each function is tightly coupled with specific source code implementations in the OpenWork repository.

Configuration Change Detection

The pool detects runtime configuration changes by computing a cryptographic fingerprint of the current engine configuration. The computeEngineConfigFingerprint function hashes the configuration state, which the system compares against currentFingerprint to determine if a rollover is necessary (source: engine-pool.ts#L215-L236).

Idle Engine In-Place Reload

When the primary engine is idle, the pool optimizes for speed by reloading the existing process with new configuration rather than spawning a fresh instance. This path is triggered through requestRollover, which checks engine busyness via the engineBusy hook before proceeding with the in-place refresh (source: engine-pool.ts#L332-L352).

Blue/Green Standby Creation

For busy engines, the pool implements a full blue/green deployment pattern. The rollOver function spawns a standby engine with the updated configuration, routes new incoming requests to this standby, and maintains the old engine in a draining state. This ensures zero-downtime transitions even when users have active sessions (source: engine-pool.ts#L560-L608).

Session Ownership and Request Pinning

To prevent request routing errors during transitions, the pool maintains strict session ownership tracking. The sessionOwnership map and pinnedRequests handling within routeRequest ensure that requests belonging to existing sessions continue to hit the correct engine generation—either the primary or the draining instance—until the session naturally concludes (source: engine-pool.ts#L124-L166).

Graceful Draining and Shutdown

The startDrainMonitor function tracks engines in the draining state, monitoring active session counts and implementing a configurable timeout mechanism. If sessions linger beyond the grace period, the system forcibly aborts them before calling retire to clean up the process (source: engine-pool.ts#L714-L782).

Recovery from Engine Crashes

When the primary engine dies unexpectedly, recoverDeadPrimary and runDeadPrimaryRecovery automatically spawn a replacement engine and promote it to primary status. This self-healing capability ensures that single-engine failures do not result in service outages (source: engine-pool.ts#L842-L904).

Metrics and Observability

Throughout the file, structured logging via hooks.logger?.log captures rollover steps, connection failures, and drain events. The snapshot() function exposes the current pool state—including generation roles, process IDs, ports, and remaining drain time—for external monitoring systems.

How the Blue/Green Rollover Works

The engine pool orchestrates configuration updates through a four-state state machine: primary, starting, draining, and dead.

When a configuration change is detected:

  1. Assessment: The system checks if the current primary is idle using the engineBusy hook.
  2. Decision: If idle, it performs an in-place reload via writeRuntimeConfigFile and postRefreshSync. If busy, it spawns a standby.
  3. Transition: New requests route to the standby (green) while existing sessions continue on the draining (blue) engine.
  4. Cleanup: Once all sessions complete or the drain timeout expires, disposeAll or retire terminates the old process.

This workflow ensures that no active requests are dropped during configuration updates, fulfilling the zero-downtime guarantee.

Public API and Key Functions

The module exports several high-level functions that the OpenWork server uses to interact with the pool:

  • enginePoolForConfig – Retrieves or initializes a pool instance for a specific server configuration.
  • setEnginePoolForConfig – Registers a new pool instance in the global registry.
  • requestRollover – Initiates the blue/green rollover process.
  • primaryUrl – Returns the HTTP endpoint of the current primary engine (source: engine-pool.ts#L104-L108).
  • snapshot – Returns a serializable state object showing all active generations and their health status (source: engine-pool.ts#L190-L207).
  • connections – Provides access to the active connection registry.
  • disposeAll – Gracefully shuts down all engines in the pool, used during server shutdown (source: engine-pool.ts#L1005-L1035).

Practical Usage Examples

Requesting a Manual Rollover

Administrators can trigger a rollover explicitly after updating model configurations:

await enginePool.requestRollover({
  reason: "User updated model list",
  workspace: currentWorkspace,
  manual: true,               // Explicit admin action
  awaitPostRefreshSync: true, // Wait for MCP sync before completing
});

Routing Requests to the Primary Engine

The HTTP server uses this method to determine where to proxy incoming requests:

const primaryUrl = enginePool.primaryUrl(); // → "http://127.0.0.1:3000"

Monitoring Pool State

Health check endpoints can expose detailed pool status:

const snapshot = enginePool.snapshot();
console.log(snapshot.generations);
/*
[
  { role: "primary", pid: 12345, port: 3000, spawnedAt: 1723939200000, drainRemainingMs: null },
  { role: "draining", pid: 12346, port: 3001, spawnedAt: 1723939250000, drainRemainingMs: 120000 }
]
*/

Graceful Shutdown

During server termination, cleanup is handled through:

await enginePool.disposeAll(); // Closes primary + any draining engines

The engine pool cooperates with several adjacent modules in the OpenWork codebase:

  • engine-registry.ts – Persists engine instances in the OpenWork registry, enabling the pool to register new engines and retire old ones.
  • managed-opencode.ts – Spawns managed OpenCode processes; the pool calls this when creating standby engines or replacement primaries.
  • server-fetch.ts – Provides low-level HTTP utilities for loopback health checks and session queries against running engines.
  • types.ts – Contains TypeScript definitions for ServerConfig, WorkspaceInfo, and other interfaces consumed by the pool's configuration fingerprinting logic.

Summary

  • engine-pool.ts implements a blue/green rollover manager that enables zero-downtime configuration updates for the OpenWork server.
  • The module detects config changes via fingerprinting and chooses between in-place reloads or standby creation based on engine busyness.
  • Session ownership tracking and request pinning ensure that active user sessions complete on the correct engine generation during transitions.
  • Automatic crash recovery via recoverDeadPrimary provides self-healing capabilities for the OpenCode engine.
  • The public API exposes functions like requestRollover, primaryUrl, and disposeAll for integration with the broader server infrastructure.

Frequently Asked Questions

What triggers a rollover in engine-pool.ts?

A rollover triggers when the computeEngineConfigFingerprint function detects a mismatch between the current runtime configuration and the active engine's fingerprint. This can occur automatically when configuration files change, or manually when an administrator calls requestRollover with the manual: true flag.

How does engine-pool.ts handle active sessions during updates?

The pool maintains a sessionOwnership map that pins requests to their originating engine generation. When a standby spawns, new sessions route to the fresh engine while existing sessions continue hitting the draining engine. The startDrainMonitor tracks these sessions and only retires the old engine once all sessions complete or the drain timeout expires.

What happens if the primary engine crashes unexpectedly?

The runDeadPrimaryRecovery function automatically detects the failure through health check polling and spawns a replacement engine. This recovery mechanism promotes the new engine to primary status within seconds, ensuring that the OpenWork server remains operational even during catastrophic engine failures.

Can engine-pool.ts run multiple engines simultaneously?

Yes, during a rollover operation, the pool intentionally runs both a primary and a draining engine concurrently. The snapshot() function can report on multiple generations simultaneously, including their respective process IDs, ports, and drain remaining times.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →