Shannon's Temporal Workflow Crash Recovery: How It Preserves State in the Pentest Pipeline
Shannon's Temporal workflow crash recovery automatically persists workflow state at every await boundary, allowing the pentest pipeline to resume exactly where it left off after worker crashes without losing progress or data.
Shannon, an open-source penetration testing automation framework by KeygraphHQ, leverages Temporal IO to ensure its multi-agent pentest pipeline survives infrastructure failures. The implementation relies on Temporal's deterministic replay mechanism, where every await point acts as a durable checkpoint that preserves the workflow's memory to an internal event store.
How Temporal Crash Recovery Works in Shannon
Durable Replay Boundaries and State Persistence
Temporal IO provides automatic crash-recovery by treating every await in a workflow as a durable replay boundary. When a workflow yields at an await point—such as when calling an activity—Temporal persists the workflow's entire memory state to its internal event store. If the worker process dies, the Temporal server reloads this state and continues execution from the last persisted point.
The State Object and Checkpointing Strategy
In src/temporal/workflows.ts, Shannon implements a centralized state object that holds all mutable workflow data:
// src/temporal/workflows.ts lines 107-118
const state: WorkflowState = {
workflowId,
status: 'running',
currentPhase: 'init',
currentAgent: null,
completedAgents: [],
metrics: {
startTime: Date.now(),
costs: {},
turnCounts: {},
},
error: null,
};
All mutations to state occur before the next await. When the workflow yields—for example, await a.runPreReconAgent(...)—Temporal writes the entire state to its event store. This ensures that if the worker crashes, the resumed workflow sees the exact same state values.
Implementation Details in the Shannon Codebase
Workflow State Management (workflows.ts)
The pentestPipelineWorkflow function in src/temporal/workflows.ts orchestrates the entire pipeline. It registers a query handler that allows external tools to read the latest persisted state at any time:
// src/temporal/workflows.ts lines 20-25
setHandler(getProgress, () => ({
...state,
elapsedMs: Date.now() - state.metrics.startTime,
completedCount: state.completedAgents.length,
}));
This query returns the durably stored state even if the worker process has restarted, demonstrating that the state persists across process boundaries.
Worker Resilience and Restart Capability (worker.ts)
The Temporal worker implementation in src/temporal/worker.ts handles activity execution. If this process crashes, the Temporal server retains the workflow's history. A new worker instance can be started without code changes and will automatically pick up pending activities:
// src/temporal/worker.ts lines 34-53
const worker = await Worker.create({
connection,
namespace: 'default',
taskQueue: 'shannon-pipeline',
workflowsPath: require.resolve('./workflows'),
activities,
maxConcurrentActivityTaskExecutions: 3,
});
await worker.run();
This architecture means that worker crashes are transparent to the workflow execution—the pipeline continues from the last checkpoint once a new worker connects.
Querying Persisted State (client.ts)
The CLI client in src/temporal/client.ts demonstrates how external processes interact with the durable state. It polls the getProgress query while the workflow runs:
// src/temporal/client.ts lines 91-100
const progressInterval = setInterval(async () => {
const progress = await handle.query('getProgress');
console.log(
`[${Math.floor(progress.elapsedMs / 1000)}s]`,
`Phase: ${progress.currentPhase || 'unknown'}`,
`Agent: ${progress.currentAgent || 'none'}`,
`Completed: ${progress.completedAgents.length}/13`
);
}, 30_000);
This polling continues to work even if the original worker crashes and a new one takes over, because the query reads from Temporal's persisted event store, not from the worker's ephemeral memory.
End-to-End Crash Recovery Flow
The complete crash recovery mechanism follows this sequence:
-
Workflow Start:
client.tsinitiatespentestPipelineWorkflowwith a uniqueworkflowIdstored in the input (lines 68-75). -
First Checkpoint: The workflow executes
await a.runPreReconAgent(...)(lines 45-48). Upon completion, Temporal persists the entirestateobject to its event store. -
Worker Crash: The process running
worker.tsexits unexpectedly. The Temporal server retains the workflow history and the last persistedstate. -
Worker Restart: A new worker instance starts (
npm run temporal:worker). It connects to the same namespace and task queue (lines 34-53). -
Execution Resumes: Temporal assigns the pending activity tasks to the new worker. The workflow resumes with the same
statevalues it had at the lastawait, continuing from thereconphase without re-executing completed agents. -
Progress Monitoring: Throughout this process,
client.tscontinues pollinggetProgress, receiving accurate state snapshots even across worker restarts (lines 93-99).
Handling Failures and Non-Retryable Errors
Shannon configures retry behavior through PRODUCTION_RETRY and TESTING_RETRY objects (lines 44-68). These define backoff strategies and specify non-retryable error types such as AuthenticationError.
When an activity fails with a non-retryable error, Temporal aborts that activity. The workflow's catch block (lines 303-327) logs the failure, records the current state with status: 'failed', and re-throws the error. Because state was persisted at the last await, operators can query the workflow to see exactly which agent failed and at what phase, enabling precise debugging and potential manual resumption.
Code Examples
Starting a Pipeline and Monitoring Progress
// src/temporal/client.ts – fragment
const handle = await client.workflow.start(
'pentestPipelineWorkflow',
{
taskQueue: 'shannon-pipeline',
workflowId,
args: [input],
}
);
// Periodic progress query (uses persisted state)
const progressInterval = setInterval(async () => {
const progress = await handle.query('getProgress');
console.log(
`[${Math.floor(progress.elapsedMs / 1000)}s]`,
`Phase: ${progress.currentPhase || 'unknown'}`,
`Agent: ${progress.currentAgent || 'none'}`,
`Completed: ${progress.completedAgents.length}/13`
);
}, 30_000);
The query reads the durably stored state even if the worker process restarts.
Simulating Worker Crash and Recovery
# 1. Start the workflow
npm run temporal:start -- https://example.com /path/to/repo --wait
# 2. In another terminal, kill the worker (simulates a crash)
pkill -f shannon-worker # or Ctrl‑C if running in foreground
# 3. Restart the worker – Temporal will pick up where it left off
npm run temporal:worker
No code changes are required; Temporal automatically reloads the workflow's persisted state and continues the pending activities.
Inspecting State After Recovery
// In any Node process with Temporal client
const progress = await client.workflow.handle(workflowId).query('getProgress');
console.log('Recovered state:', progress);
The output shows currentPhase, currentAgent, and all metrics collected up to the point of failure, enabling precise debugging.
Summary
- Durable replay boundaries: Every
awaitinsrc/temporal/workflows.tsacts as a checkpoint where Temporal persists the workflow'sstateobject to its event store. - Automatic recovery: If the worker process in
src/temporal/worker.tscrashes, a new instance can start without code changes and resume execution from the last persisted state. - State preservation: The
stateobject (lines 107-118) holds phase, agent progress, and metrics, ensuring no data loss across restarts. - Observable progress: The
getProgressquery handler (lines 20-25) exposes the durable state to external clients, enabling real-time monitoring even during crash recovery. - Configurable resilience: Retry policies in
PRODUCTION_RETRY(lines 44-68) define which errors are retryable, while the catch block (lines 303-327) ensures final state is recorded before failure.
Frequently Asked Questions
How does Shannon's Temporal workflow crash recovery handle worker process failures?
When the worker process running src/temporal/worker.ts crashes, the Temporal server retains the workflow's execution history and the last persisted state from the most recent await checkpoint. A new worker instance can be started immediately without code changes, and Temporal automatically assigns pending activities to the new worker, resuming execution from the exact point of interruption.
What specific data does Shannon preserve in the workflow state?
The state object defined in src/temporal/workflows.ts (lines 107-118) preserves the workflowId, current execution status, currentPhase (e.g., 'init', 'recon'), currentAgent being executed, an array of completedAgents, and a metrics object containing start time, costs, and turn counts. This comprehensive snapshot ensures zero data loss during crash recovery.
Can external tools monitor workflow progress during crash recovery?
Yes. Shannon registers a getProgress query handler in src/temporal/workflows.ts (lines 20-25) that returns the current state object. External clients like the CLI in src/temporal/client.ts poll this query every 30 seconds to display real-time progress. Because the query reads from Temporal's durable event store rather than the worker's memory, it returns accurate state information even when the original worker has crashed and a new one has taken over.
How does Shannon handle non-retryable errors during workflow execution?
Shannon configures retry policies through PRODUCTION_RETRY and TESTING_RETRY objects in src/temporal/workflows.ts (lines 44-68), which specify backoff strategies and nonRetryableErrorTypes such as AuthenticationError. When an activity fails with a non-retryable error, Temporal aborts that specific activity. The workflow's catch block (lines 303-327) then logs the failure, updates state.status to 'failed', and re-throws the error while preserving the final state in Temporal's store, enabling operators to inspect exactly where and why the workflow failed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →