# How Resident Workers Maintain Scheduling Across Supervisor Replacement in Prime Agent

> Learn how resident workers maintain scheduling across supervisor replacement using durable JSON descriptors and replay aware adoption protocols in Prime Agent. Ensure seamless transitions.

- Repository: [Prime Intellect/prime-agent](https://github.com/PrimeIntellect-ai/prime-agent)
- Tags: internals
- Published: 2026-09-05

---

**Resident workers preserve their cron jobs, heartbeat state, and session scheduling across supervisor replacement through durable JSON descriptors stored on disk and a replay-aware adoption protocol that allows new supervisors to reconnect to existing worker processes without interruption.**

Prime Agent's daemon architecture intentionally separates the *supervisor* process—which owns the Unix socket and coordinates execution—from *resident workers* that host persistent agent sessions. When the supervisor restarts to apply updates or recover from failures, the system must hand off control without dropping scheduled work or terminating active sessions. This deep dive into the PrimeIntellect-ai/prime-agent source code reveals the exact mechanisms that make this seamless transition possible.

## Persistent Worker State on Disk

The foundation of scheduling continuity lies in durable worker descriptors that survive process restarts. Each resident worker maintains a persistent record that the supervisor can reload and reconcile after a crash or update.

### The ResidentWorker Descriptor Structure

In [`packages/coding-agent/src/modes/daemon/daemon-supervisor.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/packages/coding-agent/src/modes/daemon/daemon-supervisor.ts) (lines [28‑44] and [46‑54]), the supervisor defines the `ResidentWorker` interface to track both runtime connections and persistent state:

```typescript
interface ResidentWorker {
  descriptor: DaemonWorkerDescriptor;
  descriptorPath: string;
  client?: DaemonWorkerClient;
  heartbeatSnapshot?: AgentConnectionHeartbeat[];
  heartbeatSnapshotStale?: boolean;
  // …
}

```

When a worker spawns, the supervisor writes a JSON descriptor to `<descriptorDir>/<workerId>.json`. This file contains the worker's socket path, session file location, and critically, the **heartbeat** and **scheduled-job** metadata (`heartbeatSnapshot`, `heartbeatSnapshotStale`). During startup, the `loadWorkerDescriptors` function (lines [308‑336]) reads every descriptor file and upgrades it to a `durableDaemonWorkerDescriptor` shape, enabling the new supervisor to reconstruct the worker's scheduling context.

## Scheduling Information Outside the Supervisor

To ensure zero-downtime hand-offs, Prime Agent stores all scheduling artifacts outside the supervisor's memory. This separation of concerns ensures that cron jobs and session state persist even when the coordinating process dies.

### Session Artifact Persistence

Scheduled jobs live as newline-delimited JSON in `session-scheduled-jobs.jsonl` files within each session directory (see [`core/cron-jobs.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/core/cron-jobs.ts)). The system writes these artifacts via `getSessionArtifactPathForFile` whenever a worker creates or updates a cron job. Because these files reside in the session's artifact directory rather than supervisor memory, they remain available for the replacement supervisor to load via `collectPassiveScheduledJobs` (lines [518‑560]).

### Heartbeat Snapshots for Liveness Detection

Workers publish periodic heartbeat messages that the supervisor accumulates in `heartbeatSnapshot` arrays within the resident worker record. These snapshots serve dual purposes: they detect worker liveness and prevent premature eviction during supervisor transitions. When a new supervisor adopts a worker, it restores the latest snapshot to maintain continuity in health monitoring.

## Supervisor Startup and Worker Adoption

When a new supervisor process launches, it executes a specific recovery sequence to assume control of resident workers without disrupting their scheduled tasks.

### Acquiring Ownership and Loading State

The adoption process begins with `acquireDaemonSupervisorOwnership` (lines [734‑777]) in [`daemon-supervisor-ownership.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-supervisor-ownership.ts), which obtains a fresh socket lease and prepares the socket path via `prepareDaemonSocketPath`. The new supervisor then invokes `loadWorkerDescriptors` (lines [308‑336]) to read all persisted worker JSON files from disk. Each loaded descriptor receives a `lifecycle: "recovering"` flag, signaling that the worker may already be running and requires reconnection rather than fresh initialization.

### Reconnecting to Existing Workers

The `adoptOrRecoverWorker` function (called from `start` at lines [891‑933]) attempts to reconnect to each worker via its `workerSocketPath`. If the worker process remains alive, the supervisor re-establishes the `DaemonWorkerClient` connection and restores both heartbeat snapshots and job state. If the worker died during the transition, the supervisor cleans up the stale descriptor and may spawn a replacement.

## Maintaining Scheduling Continuity

Once reconnected, the supervisor ensures that time-based scheduling and eviction logic continue operating from the exact state left by the previous generation.

### Cron Job Execution Persistence

Because cron jobs live as session artifacts on disk, newly adopted workers automatically reload them when initializing their sessions. The supervisor's idle-eviction sweep (`runIdleEvictionSweep`) and scheduled-wake recompute logic (`recomputeScheduledSessionWake`) both read from these artifact files, ensuring that jobs fire on schedule even after a supervisor restart. The system never loses track of pending executions because the canonical state lives in the filesystem, not volatile memory.

### Heartbeat State Restoration

The `workerEvictionSnapshot` logic (lines [610‑664]) relies on the `heartbeatSnapshot` persisted in the worker descriptor. When the new supervisor adopts a worker, it restores this snapshot to prevent the eviction timer from resetting to zero, which would otherwise risk terminating active workers immediately after a hand-off.

## Replacement-Specific Hand-Off Protocol

When the supervisor initiates a controlled restart—such as during an update—it follows a graceful replacement protocol to minimize scheduling disruption.

### Draining and Signaling

The supervisor enters a drain phase using `withEvictionFence` to pause mutations, then signals workers via `prepare_update_restart` to stop accepting new commands. This ensures scheduling state stabilizes before the socket transfers to the new process.

### Replacement Connections and Catch-Up

Workers may open special *replacement* connections identified by `snapshotPurpose: "replacement"` in [`daemon-worker-protocol.ts`](https://github.com/PrimeIntellect-ai/prime-agent/blob/main/daemon-worker-protocol.ts) (lines [53‑56]). These connections allow the new supervisor to attach without losing in-flight data. The new supervisor receives a replacement snapshot via `queueCatchup(..., "replacement")` (lines [572‑585]) and merges it with the existing state, preserving scheduled jobs and heartbeat sequences across the transition.

## Practical Code Examples

### Inspecting Worker Scheduling State

To verify that a worker's scheduling information persists correctly, you can inspect its descriptor file directly:

```typescript
import { readFileSync } from "node:fs";
import { join } from "node:path";

const descriptorDir = "/path/to/agent/daemon-workers";
const workerId = "abc123";
const descriptorPath = join(descriptorDir, `${workerId}.json`);
const descriptor = JSON.parse(readFileSync(descriptorPath, "utf8"));

console.log("Worker heartbeat snapshot:", descriptor.heartbeatSnapshot);
console.log("Scheduled jobs stored in:", descriptor.sessionFile);
console.log("Last cron update:", descriptor.lastCronUpdateTimestamp);

```

### Simulating a Supervisor Replacement

You can test the replacement protocol by spawning a new supervisor process that takes over the socket:

```typescript
import { spawn } from "node:child_process";

const newSupervisorSock = "/tmp/prime-agent-supervisor-new.sock";

// Start a fresh supervisor that will adopt existing workers
const newSupervisor = spawn("node", [
  "dist/packages/coding-agent/src/modes/daemon/daemon-supervisor.js",
  "--socket", newSupervisorSock,
], {
  env: { 
    ...process.env, 
    PRIME_AGENT_INTERNAL_DAEMON_SUPERVISOR_REGISTRY_DIR: "/tmp/supervisor-registry" 
  },
});

newSupervisor.stdout.on("data", (data) => {
  console.log(`Supervisor: ${data}`);
});

```

When this new supervisor starts, it will execute `loadWorkerDescriptors` and `adoptOrRecoverWorker` to reconnect to resident workers, allowing their scheduled cron jobs to continue firing without missing a beat.

## Summary

- **Worker descriptors** are persisted to JSON files on disk, containing `heartbeatSnapshot` and scheduling metadata that survive supervisor crashes.
- **Session artifacts** store cron jobs as `session-scheduled-jobs.jsonl` files, ensuring scheduling state exists outside the supervisor's memory.
- **Adoption protocol** uses `acquireDaemonSupervisorOwnership`, `loadWorkerDescriptors`, and `adoptOrRecoverWorker` to reconnect to existing workers without restarting them.
- **Replacement connections** with `snapshotPurpose: "replacement"` and `queueCatchup` enable graceful hand-offs during controlled supervisor updates.
- **Eviction logic** relies on restored heartbeat snapshots to avoid terminating workers immediately after a supervisor transition.

## Frequently Asked Questions

### What happens to scheduled cron jobs when the Prime Agent supervisor restarts?

Scheduled cron jobs survive supervisor restarts because they are stored as session artifacts in `session-scheduled-jobs.jsonl` files on disk rather than in supervisor memory. When the new supervisor loads worker descriptors via `loadWorkerDescriptors`, it reconstructs the scheduling state from these files. The `recomputeScheduledSessionWake` function then ensures that cron jobs fire at their correct times without requiring any action from the newly restarted workers.

### How does the new supervisor know which workers are still active?

The supervisor determines worker liveness through the `adoptOrRecoverWorker` function (lines [891‑933]), which attempts to connect via the `workerSocketPath` stored in each worker's JSON descriptor. If the worker responds, the supervisor restores the connection and updates its internal `DaemonWorkerClient` mapping. If the connection fails, the supervisor treats the worker as dead and cleans up the stale descriptor. The `heartbeatSnapshot` stored in the descriptor provides additional context about when the worker was last seen alive.

### Where is the worker's scheduling state stored during a supervisor hand-off?

Worker scheduling state exists in three locations: the **worker descriptor** JSON file (containing `heartbeatSnapshot` and metadata), the **session artifacts** directory (containing `session-scheduled-jobs.jsonl`), and the **replacement connection** buffer (containing in-flight snapshots during controlled restarts). This redundant storage ensures that even if the supervisor crashes mid-hand-off, the new instance can fully reconstruct the worker's scheduling context from disk.

### Can resident workers continue executing tasks while the supervisor is being replaced?

Yes. Resident workers operate as independent processes hosting agent sessions. During a controlled replacement, the supervisor signals workers to enter `prepare_update_restart` state, which stops new command acceptance but allows ongoing tasks to complete. Workers maintain their own cron job schedules by reading from session artifacts, and the new supervisor reconnects via `snapshotPurpose: "replacement"` connections to catch up on any state changes that occurred during the transition.