# Shannon's Temporal Workflow Crash Recovery: How It Preserves State in the Pentest Pipeline

> Discover Shannon's Temporal workflow crash recovery. Learn how it automatically persists state at await boundaries to resume pentest pipelines seamlessly after worker crashes, preserving all progress and data.

- Repository: [KeygraphHQ/shannon](https://github.com/keygraphhq/shannon)
- Tags: internals
- Published: 2026-02-16

---

**Shannon's Temporal workflow crash recovery automatically persists workflow state at every await boundary, allowing the pentest pipeline to resume exactly where it left off after worker crashes without losing progress or data.**

Shannon, an open-source penetration testing automation framework by KeygraphHQ, leverages Temporal IO to ensure its multi-agent pentest pipeline survives infrastructure failures. The implementation relies on Temporal's deterministic replay mechanism, where every `await` point acts as a durable checkpoint that preserves the workflow's memory to an internal event store.

## How Temporal Crash Recovery Works in Shannon

### Durable Replay Boundaries and State Persistence

Temporal IO provides **automatic crash-recovery** by treating every `await` in a workflow as a **durable replay boundary**. When a workflow yields at an await point—such as when calling an activity—Temporal persists the workflow's entire memory state to its internal event store. If the worker process dies, the Temporal server reloads this state and continues execution from the last persisted point.

### The State Object and Checkpointing Strategy

In [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts), Shannon implements a centralized `state` object that holds all mutable workflow data:

```typescript
// src/temporal/workflows.ts lines 107-118
const state: WorkflowState = {
  workflowId,
  status: 'running',
  currentPhase: 'init',
  currentAgent: null,
  completedAgents: [],
  metrics: {
    startTime: Date.now(),
    costs: {},
    turnCounts: {},
  },
  error: null,
};

```

All mutations to `state` occur **before** the next `await`. When the workflow yields—for example, `await a.runPreReconAgent(...)`—Temporal writes the entire `state` to its event store. This ensures that if the worker crashes, the resumed workflow sees the exact same `state` values.

## Implementation Details in the Shannon Codebase

### Workflow State Management (workflows.ts)

The `pentestPipelineWorkflow` function in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) orchestrates the entire pipeline. It registers a **query handler** that allows external tools to read the latest persisted state at any time:

```typescript
// src/temporal/workflows.ts lines 20-25
setHandler(getProgress, () => ({
  ...state,
  elapsedMs: Date.now() - state.metrics.startTime,
  completedCount: state.completedAgents.length,
}));

```

This query returns the **durably stored** `state` even if the worker process has restarted, demonstrating that the state persists across process boundaries.

### Worker Resilience and Restart Capability (worker.ts)

The Temporal worker implementation in [`src/temporal/worker.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/worker.ts) handles activity execution. If this process crashes, the Temporal server retains the workflow's history. A new worker instance can be started without code changes and will automatically pick up pending activities:

```typescript
// src/temporal/worker.ts lines 34-53
const worker = await Worker.create({
  connection,
  namespace: 'default',
  taskQueue: 'shannon-pipeline',
  workflowsPath: require.resolve('./workflows'),
  activities,
  maxConcurrentActivityTaskExecutions: 3,
});

await worker.run();

```

This architecture means that **worker crashes are transparent** to the workflow execution—the pipeline continues from the last checkpoint once a new worker connects.

### Querying Persisted State (client.ts)

The CLI client in [`src/temporal/client.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/client.ts) demonstrates how external processes interact with the durable state. It polls the `getProgress` query while the workflow runs:

```typescript
// src/temporal/client.ts lines 91-100
const progressInterval = setInterval(async () => {
  const progress = await handle.query('getProgress');
  console.log(
    `[${Math.floor(progress.elapsedMs / 1000)}s]`,
    `Phase: ${progress.currentPhase || 'unknown'}`,
    `Agent: ${progress.currentAgent || 'none'}`,
    `Completed: ${progress.completedAgents.length}/13`
  );
}, 30_000);

```

This polling continues to work even if the original worker crashes and a new one takes over, because the query reads from Temporal's persisted event store, not from the worker's ephemeral memory.

## End-to-End Crash Recovery Flow

The complete crash recovery mechanism follows this sequence:

1. **Workflow Start**: [`client.ts`](https://github.com/KeygraphHQ/shannon/blob/main/client.ts) initiates `pentestPipelineWorkflow` with a unique `workflowId` stored in the input (lines 68-75).

2. **First Checkpoint**: The workflow executes `await a.runPreReconAgent(...)` (lines 45-48). Upon completion, Temporal persists the entire `state` object to its event store.

3. **Worker Crash**: The process running [`worker.ts`](https://github.com/KeygraphHQ/shannon/blob/main/worker.ts) exits unexpectedly. The Temporal server retains the workflow history and the last persisted `state`.

4. **Worker Restart**: A new worker instance starts (`npm run temporal:worker`). It connects to the same namespace and task queue (lines 34-53).

5. **Execution Resumes**: Temporal assigns the pending activity tasks to the new worker. The workflow resumes with the same `state` values it had at the last `await`, continuing from the `recon` phase without re-executing completed agents.

6. **Progress Monitoring**: Throughout this process, [`client.ts`](https://github.com/KeygraphHQ/shannon/blob/main/client.ts) continues polling `getProgress`, receiving accurate state snapshots even across worker restarts (lines 93-99).

## Handling Failures and Non-Retryable Errors

Shannon configures retry behavior through `PRODUCTION_RETRY` and `TESTING_RETRY` objects (lines 44-68). These define backoff strategies and specify **non-retryable error types** such as `AuthenticationError`.

When an activity fails with a non-retryable error, Temporal aborts that activity. The workflow's `catch` block (lines 303-327) logs the failure, records the current `state` with `status: 'failed'`, and re-throws the error. Because `state` was persisted at the last `await`, operators can query the workflow to see exactly which agent failed and at what phase, enabling precise debugging and potential manual resumption.

## Code Examples

### Starting a Pipeline and Monitoring Progress

```typescript
// src/temporal/client.ts – fragment
const handle = await client.workflow.start(
  'pentestPipelineWorkflow',
  {
    taskQueue: 'shannon-pipeline',
    workflowId,
    args: [input],
  }
);

// Periodic progress query (uses persisted state)
const progressInterval = setInterval(async () => {
  const progress = await handle.query('getProgress');
  console.log(
    `[${Math.floor(progress.elapsedMs / 1000)}s]`,
    `Phase: ${progress.currentPhase || 'unknown'}`,
    `Agent: ${progress.currentAgent || 'none'}`,
    `Completed: ${progress.completedAgents.length}/13`
  );
}, 30_000);

```

The query reads the **durably stored** `state` even if the worker process restarts.

### Simulating Worker Crash and Recovery

```bash

# 1. Start the workflow

npm run temporal:start -- https://example.com /path/to/repo --wait

# 2. In another terminal, kill the worker (simulates a crash)

pkill -f shannon-worker   # or Ctrl‑C if running in foreground

# 3. Restart the worker – Temporal will pick up where it left off

npm run temporal:worker

```

No code changes are required; Temporal automatically reloads the workflow's persisted state and continues the pending activities.

### Inspecting State After Recovery

```typescript
// In any Node process with Temporal client
const progress = await client.workflow.handle(workflowId).query('getProgress');
console.log('Recovered state:', progress);

```

The output shows `currentPhase`, `currentAgent`, and all metrics collected up to the point of failure, enabling precise debugging.

## Summary

- **Durable replay boundaries**: Every `await` in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) acts as a checkpoint where Temporal persists the workflow's `state` object to its event store.
- **Automatic recovery**: If the worker process in [`src/temporal/worker.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/worker.ts) crashes, a new instance can start without code changes and resume execution from the last persisted state.
- **State preservation**: The `state` object (lines 107-118) holds phase, agent progress, and metrics, ensuring no data loss across restarts.
- **Observable progress**: The `getProgress` query handler (lines 20-25) exposes the durable state to external clients, enabling real-time monitoring even during crash recovery.
- **Configurable resilience**: Retry policies in `PRODUCTION_RETRY` (lines 44-68) define which errors are retryable, while the catch block (lines 303-327) ensures final state is recorded before failure.

## Frequently Asked Questions

### How does Shannon's Temporal workflow crash recovery handle worker process failures?

When the worker process running [`src/temporal/worker.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/worker.ts) crashes, the Temporal server retains the workflow's execution history and the last persisted `state` from the most recent `await` checkpoint. A new worker instance can be started immediately without code changes, and Temporal automatically assigns pending activities to the new worker, resuming execution from the exact point of interruption.

### What specific data does Shannon preserve in the workflow state?

The `state` object defined in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) (lines 107-118) preserves the `workflowId`, current execution `status`, `currentPhase` (e.g., 'init', 'recon'), `currentAgent` being executed, an array of `completedAgents`, and a `metrics` object containing start time, costs, and turn counts. This comprehensive snapshot ensures zero data loss during crash recovery.

### Can external tools monitor workflow progress during crash recovery?

Yes. Shannon registers a `getProgress` query handler in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) (lines 20-25) that returns the current `state` object. External clients like the CLI in [`src/temporal/client.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/client.ts) poll this query every 30 seconds to display real-time progress. Because the query reads from Temporal's durable event store rather than the worker's memory, it returns accurate state information even when the original worker has crashed and a new one has taken over.

### How does Shannon handle non-retryable errors during workflow execution?

Shannon configures retry policies through `PRODUCTION_RETRY` and `TESTING_RETRY` objects in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) (lines 44-68), which specify backoff strategies and `nonRetryableErrorTypes` such as `AuthenticationError`. When an activity fails with a non-retryable error, Temporal aborts that specific activity. The workflow's catch block (lines 303-327) then logs the failure, updates `state.status` to `'failed'`, and re-throws the error while preserving the final state in Temporal's store, enabling operators to inspect exactly where and why the workflow failed.