# How Shannon Handles Graceful Failure of Agents in Parallel Execution Groups

> Learn how Shannon ensures successful pentests with graceful failure handling for agents in parallel execution groups using Promise.allSettled and structured error classification.

- Repository: [KeygraphHQ/shannon](https://github.com/keygraphhq/shannon)
- Tags: internals
- Published: 2026-02-16

---

**Shannon uses `Promise.allSettled` combined with structured error classification to ensure that when individual vulnerability or exploit agents fail, the remaining parallel pipelines continue executing and the overall pentest workflow completes successfully.**

The KeygraphHQ/shannon repository implements a robust concurrency model for automated security testing. When running multiple vulnerability-to-exploit pipelines simultaneously, graceful failure of agents in parallel execution groups prevents a single faulty agent from aborting the entire assessment. This architecture isolates failures, logs them for analysis, and allows successful agents to contribute their findings to the final report.

## The Core Mechanism: Promise.allSettled and Structured Error Handling

Shannon's resilience stems from two complementary strategies implemented in its Temporal workflows and activities.

### Parallel Execution Without Cascading Failures

In [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) at line 198, Shannon launches five concurrent pipelines (injection, XSS, authentication, SSRF, and authorization) using `Promise.allSettled` instead of `Promise.all`:

```typescript
const pipelineResults = await Promise.allSettled([
  runVulnExploitPipeline('injection', () => a.runInjectionVulnAgent(activityInput), () => a.runInjectionExploitAgent(activityInput)),
  runVulnExploitPipeline('xss',       () => a.runXssVulnAgent(activityInput),       () => a.runXssExploitAgent(activityInput)),
  runVulnExploitPipeline('auth',      () => a.runAuthVulnAgent(activityInput),      () => a.runAuthExploitAgent(activityInput)),
  runVulnExploitPipeline('ssrf',      () => a.runSsrfVulnAgent(activityInput),      () => a.runSsrfExploitAgent(activityInput)),
  runVulnExploitPipeline('authz',     () => a.runAuthzVulnAgent(activityInput),     () => a.runAuthzExploitAgent(activityInput)),
]);

```

Unlike `Promise.all`, which rejects immediately when any promise fails, `Promise.allSettled` waits for every pipeline to complete. It returns an array of objects indicating whether each pipeline fulfilled or rejected, ensuring that a failure in the XSS pipeline does not prevent the injection or SSRF pipelines from completing their work.

### Distinguishing Retryable from Non-Retryable Errors

In [`src/temporal/activities.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/activities.ts) at line 40, Shannon implements `checkExploitationQueue` to classify errors before they can disrupt the workflow:

```typescript
if (error?.retryable) {
  // Temporal will retry the whole vulnerability agent
  throw error;
}
// Non‑retryable → skip exploitation gracefully
console.log(chalk.yellow(`⚠️ ${vulnType}: ${error?.message ?? 'Unknown error'}, skipping exploitation`));
return { shouldExploit: false, shouldRetry: false, vulnerabilityCount: 0, vulnType };

```

**Retryable errors** (such as transient network timeouts) are re-thrown to trigger Temporal's built-in retry mechanism. **Non-retryable errors** (such as malformed queue files or configuration issues) are caught, logged, and converted into a skip decision. This prevents a single agent's fatal error from bubbling up and terminating the entire parallel execution group.

## Implementing Graceful Failure in Shannon's Workflow

The graceful failure strategy operates across two distinct phases: result aggregation and conditional exploitation.

### Aggregating Results from Parallel Pipelines

After `Promise.allSettled` returns, the workflow in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) processes each result independently:

```typescript
const failedPipelines: string[] = [];
for (const result of pipelineResults) {
  if (result.status === 'fulfilled') {
    const { vulnType, vulnMetrics, exploitMetrics } = result.value;
    // store successful metrics …
  } else {
    const errorMsg = result.reason instanceof Error
      ? result.reason.message
      : String(result.reason);
    failedPipelines.push(errorMsg);          // remember but keep going
  }
}
if (failedPipelines.length) {
  console.log(`⚠️ ${failedPipelines.length} pipeline(s) failed:`, failedPipelines);
}

```

Successful pipelines contribute their vulnerability and exploit metrics to the final report. Failed pipelines are logged for debugging but do not trigger workflow termination. This ensures that partial results from healthy agents are preserved even when other agents encounter critical errors.

### Skipping Exploitation When Agents Cannot Run

The `runVulnExploitPipeline` function respects the skip decision returned by `checkExploitationQueue`. When `shouldExploit` is `false`, the workflow bypasses the exploit agent entirely:

```typescript
// Conceptual flow within runVulnExploitPipeline
const decision = await checkExploitationQueue(input, vulnType);
if (!decision.shouldExploit) {
  // Skip exploitation phase for this vulnerability type
  return { vulnType, vulnMetrics: decision.vulnerabilityCount, exploitMetrics: null };
}
// Otherwise, proceed with exploitation
const exploitResult = await runExploitAgent();

```

This conditional execution ensures that corrupted or invalid vulnerability findings do not trigger unnecessary exploit attempts, while allowing other parallel pipelines to proceed unaffected.

## End-to-End Execution Flow

Shannon's graceful failure architecture follows a predictable sequence:

1. **Parallel Launch**: All five vulnerability agents (injection, XSS, auth, SSRF, authz) execute simultaneously via `Promise.allSettled` in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts).

2. **Individual Completion**: Each agent either returns vulnerability metrics or rejects with an error. The workflow waits for all agents to settle before proceeding.

3. **Exploitation Decision**: For each successful vulnerability agent, `checkExploitationQueue` in [`src/temporal/activities.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/activities.ts) validates the queue. Non-retryable errors result in a skip decision rather than a workflow failure.

4. **Conditional Exploitation**: Only agents with valid queues proceed to the exploitation phase. Failed or skipped agents contribute zero exploit metrics but do not block other pipelines.

5. **Aggregation**: The workflow collects all successful metrics, logs specific pipeline failures, and transitions to the reporting phase with whatever data was successfully gathered.

This design ensures that a single flaky agent, network timeout, or malformed configuration cannot abort a multi-hour pentest assessment, providing resilient automation for security testing workflows.

## Summary

- **Promise.allSettled** in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) executes vulnerability and exploit pipelines in parallel without allowing individual failures to cancel the entire group.
- **Structured error classification** in [`src/temporal/activities.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/activities.ts) distinguishes retryable errors (which trigger Temporal retries) from non-retryable errors (which result in graceful skips).
- **Conditional execution** ensures that exploit agents only run when `checkExploitationQueue` returns `shouldExploit: true`, preventing invalid or corrupted vulnerability data from causing downstream failures.
- **Result aggregation** collects partial successes from healthy agents while logging specific failures, ensuring the workflow produces actionable reports even when some agents fail completely.

## Frequently Asked Questions

### What happens if one vulnerability agent crashes in Shannon?

If a single vulnerability agent crashes, `Promise.allSettled` in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts) captures the rejection while allowing the other four parallel pipelines to continue executing. The workflow logs the specific error message but does not terminate, ensuring that findings from healthy agents are preserved in the final report.

### How does Shannon decide whether to retry a failed activity or skip it?

Shannon's `checkExploitationQueue` function in [`src/temporal/activities.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/activities.ts) inspects the `retryable` property of caught errors. If `error.retryable` is true, the function re-throws the error to trigger Temporal's built-in retry mechanism. If false, it returns a skip decision (`shouldExploit: false`), allowing the workflow to bypass the failed agent without aborting the entire execution group.

### Can Shannon produce a report if all exploit agents fail?

Yes. Shannon aggregates metrics from any successful vulnerability agents even if all corresponding exploit agents fail or are skipped. The reporting phase in the Temporal workflow processes whatever data was successfully gathered, producing a partial report that includes vulnerability findings without exploitation results. This ensures that valuable reconnaissance data is never lost due to downstream exploitation failures.

### Where is the parallel execution logic implemented in the Shannon codebase?

The parallel execution logic resides primarily in [`src/temporal/workflows.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/workflows.ts), specifically around line 198 where `Promise.allSettled` orchestrates the five vulnerability-to-exploit pipelines. The supporting error classification logic that enables graceful failure handling is implemented in [`src/temporal/activities.ts`](https://github.com/KeygraphHQ/shannon/blob/main/src/temporal/activities.ts) within the `checkExploitationQueue` function starting at line 40.