# How to Manage AI Agent Lifecycles with the OpenWork Engine Pool

> Manage AI agent lifecycles effortlessly with OpenWork’s Engine Pool. Simplify startup, monitoring, and shutdown to treat your AI agent as a simple, always-available endpoint.

- Repository: [Different AI/openwork](https://github.com/different-ai/openwork)
- Tags: how-to-guide
- Published: 2026-08-21

---

**OpenWork’s Engine Pool abstracts the entire AI agent lifecycle—startup, health monitoring, graceful configuration rollovers, and shutdown—so developers can treat the OpenCode engine as a simple, always-available endpoint.**

The `different-ai/openwork` repository implements a sophisticated process pool architecture to manage AI agent instances. By handling complex state transitions behind a unified interface, the **Engine Pool** ensures that configuration updates never interrupt active sessions and that crashed engines recover automatically without manual intervention.

## Core Concepts of the Engine Pool

Understanding how the pool manages state requires familiarity with its core abstractions, each implemented in [`apps/server/src/engine-pool.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/engine-pool.ts).

### The EnginePool Class

The **EnginePool class** serves as the central orchestrator. It maintains one *primary* engine that handles all traffic and, when necessary, a *draining* standby during configuration transitions. The class implements blue/green rollover logic, continuous health monitoring, and automatic recovery dead engines (lines 297-343).

### Generations and State Tracking

A **Generation** represents a concrete engine instance bundled with its metadata—including ID, status, fingerprint, and registry ID. This abstraction enables the pool to map user sessions to specific engine processes and ensures that a draining engine retires only after all its active sessions complete (lines 115-139).

### Rollover Requests and Blue/Green Deployment

When configuration changes occur—such as adding new plugins or MCP connections—the pool receives a **rollover request**. The `requestRollover()` method (lines 1022-1071) evaluates whether to reload the engine in-place (when idle) or spawn a standby that becomes the new primary while the old engine drains.

### Request Routing

Every incoming API call flows through `routeRequest()` (lines 408-440). This method selects the appropriate engine—**primary**, **draining**, or **fallback**—based on session ownership. Requests belonging to sessions started on a draining engine continue routing to that engine until the session ends, ensuring consistency during transitions.

### Automatic Recovery

If the primary engine crashes, the pool triggers **recovery** logic via `recoverDeadPrimary()` and `runDeadPrimaryRecovery()` (lines 1214-1249). After detecting three consecutive connection failures, the pool automatically spawns a replacement process and promotes it to primary without dropping incoming requests.

### Diagnostic Snapshots

The `snapshot()` method (lines 494-505) exposes a lightweight view of the pool’s current state, including active generations, assigned ports, and drain timers. This enables real-time diagnostics and UI monitoring without disrupting engine operations.

## AI Agent Lifecycle Flow

The Engine Pool manages AI agents through five distinct phases, ensuring zero-downtime updates and fault tolerance.

### 1. Initialization and Startup

At server startup, the application calls `createEnginePoolForConfig()` from [`server.ts`](https://github.com/different-ai/openwork/blob/main/server.ts) (lines 10-15) to instantiate the pool. The system then spawns the initial OpenCode process and registers it as the primary engine via `adoptPrimary()`.

### 2. Configuration Changes

When `ServerConfig` updates, the server invokes `requestRollover()` with a description of the change. The pool decides between two paths:
- **In-place reload**: Occurs when the engine is idle, reusing the existing process.
- **Blue/green rollover**: Spawns a standby engine, switches traffic to the new primary, and marks the old engine as **draining**.

### 3. Graceful Draining

While an engine drains, the pool polls it for active sessions using `nonIdleSessionIds()`. The engine continues serving existing sessions but rejects new ones. Once all sessions finish or the drain timeout expires, the pool calls `retire()` to terminate the process cleanly.

### 4. Health Monitoring and Recovery

The pool continuously tracks connection health. Upon detecting failure, it invokes `recoverDeadPrimary()` to spawn a fresh engine and flip it to primary status, ensuring high availability without manual intervention.

### 5. Shutdown

During server shutdown, `disposeAll()` (lines 1194-1214) initiates clean closure of all generations. The method waits for draining engines to finish their sessions before terminating processes, preventing data loss or interrupted operations.

## Practical Implementation Examples

The following TypeScript snippets demonstrate common interactions with the Engine Pool API.

### Creating and Initializing the Pool

Instantiate the pool with configuration templates and lifecycle hooks, then adopt the first engine:

```typescript
import { EnginePool } from "./engine-pool.js";
import { createManagedOpencodeServer } from "./managed-opencode.js";

const template = {
  cwd: "/path/to/project",
  runtimeConfigPath: "/path/to/opencode.config.json",
  env: process.env,
  reservedPorts: () => [process.env.OPENWORK_PORT ?? 5178],
};

const pool = new EnginePool({
  config,
  template,
  hooks: {
    reloadInPlace: async () => {/* reload logic */},
    engineBusy: async () => false,
    postRefreshSync: async () => {/* sync MCPS */},
    writeRuntimeConfigFile: async () => ({ path: template.runtimeConfigPath }),
    registerTrusted: () => {},
    clearTrusted: () => {},
    logger: { log: console.log },
  },
});

// Spawn and adopt the initial engine
const handle = await pool.spawn();
await pool.waitForHealthy(handle);
pool.adoptPrimary({
  handle,
  fingerprint: await computeEngineConfigFingerprint(template),
  registryId: null,
  trustedIdentity: null,
});

```

### Handling Configuration Changes with Rollover

Trigger a zero-downtime configuration update using `requestRollover()`:

```typescript
await pool.requestRollover({
  reason: "Added new plugin xyz",
  workspace: { id: "default", /* …other fields… */ },
  manual: true,
  awaitPostRefreshSync: true,
});

```

### Routing Requests to the Correct Engine

Direct incoming HTTP traffic to the appropriate engine generation:

```typescript
function handleProxy(req: Request) {
  const route = pool.routeRequest(req.method, req.url);
  if (!route) throw new Error("No engine available");

  const target = route.target;
  // Forward request using target.baseUrl, target.username, etc.
}

```

### Monitoring Pool Health with Snapshots

Capture current state for diagnostics or dashboard display:

```typescript
const snap = pool.snapshot();
console.log("Engine pool state:", JSON.stringify(snap, null, 2));

```

### Graceful Shutdown Procedures

Terminate all engine processes cleanly during application shutdown:

```typescript
await pool.disposeAll();

```

## Summary

- The **EnginePool class** in [`apps/server/src/engine-pool.ts`](https://github.com/different-ai/openwork/blob/main/apps/server/src/engine-pool.ts) orchestrates primary and draining engine instances, abstracting process management complexity.
- **Generations** track engine metadata and session ownership, enabling safe blue/green deployments without dropping active connections.
- **Rollover requests** handle configuration changes through either in-place reloads or standby promotion, managed by `requestRollover()` (lines 1022-1071).
- **Automatic recovery** monitors engine health and spawns replacements via `recoverDeadPrimary()` when crashes occur.
- **Request routing** guarantees session affinity, ensuring requests complete on the engine instance where they started, even during drains.
- The lifecycle spans five phases: initialization, configuration change, draining, health recovery, and graceful shutdown via `disposeAll()`.

## Frequently Asked Questions

### What happens to active sessions during a configuration rollover?

Active sessions continue on the **draining** engine while new traffic routes to the **primary**. The pool polls the draining engine via `nonIdleSessionIds()` and retires it only after all sessions complete or the timeout expires. This ensures zero-interruption updates for in-flight operations.

### How does the Engine Pool detect and recover from crashed engines?

The pool tracks consecutive connection failures. After three failures, it triggers `recoverDeadPrimary()` (lines 1214-1249) to spawn a replacement process and promote it to primary. This self-healing mechanism prevents prolonged outages without requiring manual restarts.

### What is the difference between primary and draining engines?

The **primary** engine receives all new requests and handles active processing. A **draining** engine is a former primary that continues serving its existing sessions but rejects new ones, allowing graceful transition during configuration updates or before retirement.

### Can I reload an engine in-place instead of spawning a new process?

Yes. The `requestRollover()` logic checks engine occupancy via the `engineBusy` hook. If the engine reports idle status, the pool calls `reloadInPlace` to refresh configuration without spawning a standby, minimizing resource usage during simple updates.