How Shannon Handles Graceful Failure of Agents in Parallel Execution Groups
Shannon uses Promise.allSettled combined with structured error classification to ensure that when individual vulnerability or exploit agents fail, the remaining parallel pipelines continue executing and the overall pentest workflow completes successfully.
The KeygraphHQ/shannon repository implements a robust concurrency model for automated security testing. When running multiple vulnerability-to-exploit pipelines simultaneously, graceful failure of agents in parallel execution groups prevents a single faulty agent from aborting the entire assessment. This architecture isolates failures, logs them for analysis, and allows successful agents to contribute their findings to the final report.
The Core Mechanism: Promise.allSettled and Structured Error Handling
Shannon's resilience stems from two complementary strategies implemented in its Temporal workflows and activities.
Parallel Execution Without Cascading Failures
In src/temporal/workflows.ts at line 198, Shannon launches five concurrent pipelines (injection, XSS, authentication, SSRF, and authorization) using Promise.allSettled instead of Promise.all:
const pipelineResults = await Promise.allSettled([
runVulnExploitPipeline('injection', () => a.runInjectionVulnAgent(activityInput), () => a.runInjectionExploitAgent(activityInput)),
runVulnExploitPipeline('xss', () => a.runXssVulnAgent(activityInput), () => a.runXssExploitAgent(activityInput)),
runVulnExploitPipeline('auth', () => a.runAuthVulnAgent(activityInput), () => a.runAuthExploitAgent(activityInput)),
runVulnExploitPipeline('ssrf', () => a.runSsrfVulnAgent(activityInput), () => a.runSsrfExploitAgent(activityInput)),
runVulnExploitPipeline('authz', () => a.runAuthzVulnAgent(activityInput), () => a.runAuthzExploitAgent(activityInput)),
]);
Unlike Promise.all, which rejects immediately when any promise fails, Promise.allSettled waits for every pipeline to complete. It returns an array of objects indicating whether each pipeline fulfilled or rejected, ensuring that a failure in the XSS pipeline does not prevent the injection or SSRF pipelines from completing their work.
Distinguishing Retryable from Non-Retryable Errors
In src/temporal/activities.ts at line 40, Shannon implements checkExploitationQueue to classify errors before they can disrupt the workflow:
if (error?.retryable) {
// Temporal will retry the whole vulnerability agent
throw error;
}
// Non‑retryable → skip exploitation gracefully
console.log(chalk.yellow(`⚠️ ${vulnType}: ${error?.message ?? 'Unknown error'}, skipping exploitation`));
return { shouldExploit: false, shouldRetry: false, vulnerabilityCount: 0, vulnType };
Retryable errors (such as transient network timeouts) are re-thrown to trigger Temporal's built-in retry mechanism. Non-retryable errors (such as malformed queue files or configuration issues) are caught, logged, and converted into a skip decision. This prevents a single agent's fatal error from bubbling up and terminating the entire parallel execution group.
Implementing Graceful Failure in Shannon's Workflow
The graceful failure strategy operates across two distinct phases: result aggregation and conditional exploitation.
Aggregating Results from Parallel Pipelines
After Promise.allSettled returns, the workflow in src/temporal/workflows.ts processes each result independently:
const failedPipelines: string[] = [];
for (const result of pipelineResults) {
if (result.status === 'fulfilled') {
const { vulnType, vulnMetrics, exploitMetrics } = result.value;
// store successful metrics …
} else {
const errorMsg = result.reason instanceof Error
? result.reason.message
: String(result.reason);
failedPipelines.push(errorMsg); // remember but keep going
}
}
if (failedPipelines.length) {
console.log(`⚠️ ${failedPipelines.length} pipeline(s) failed:`, failedPipelines);
}
Successful pipelines contribute their vulnerability and exploit metrics to the final report. Failed pipelines are logged for debugging but do not trigger workflow termination. This ensures that partial results from healthy agents are preserved even when other agents encounter critical errors.
Skipping Exploitation When Agents Cannot Run
The runVulnExploitPipeline function respects the skip decision returned by checkExploitationQueue. When shouldExploit is false, the workflow bypasses the exploit agent entirely:
// Conceptual flow within runVulnExploitPipeline
const decision = await checkExploitationQueue(input, vulnType);
if (!decision.shouldExploit) {
// Skip exploitation phase for this vulnerability type
return { vulnType, vulnMetrics: decision.vulnerabilityCount, exploitMetrics: null };
}
// Otherwise, proceed with exploitation
const exploitResult = await runExploitAgent();
This conditional execution ensures that corrupted or invalid vulnerability findings do not trigger unnecessary exploit attempts, while allowing other parallel pipelines to proceed unaffected.
End-to-End Execution Flow
Shannon's graceful failure architecture follows a predictable sequence:
-
Parallel Launch: All five vulnerability agents (injection, XSS, auth, SSRF, authz) execute simultaneously via
Promise.allSettledinsrc/temporal/workflows.ts. -
Individual Completion: Each agent either returns vulnerability metrics or rejects with an error. The workflow waits for all agents to settle before proceeding.
-
Exploitation Decision: For each successful vulnerability agent,
checkExploitationQueueinsrc/temporal/activities.tsvalidates the queue. Non-retryable errors result in a skip decision rather than a workflow failure. -
Conditional Exploitation: Only agents with valid queues proceed to the exploitation phase. Failed or skipped agents contribute zero exploit metrics but do not block other pipelines.
-
Aggregation: The workflow collects all successful metrics, logs specific pipeline failures, and transitions to the reporting phase with whatever data was successfully gathered.
This design ensures that a single flaky agent, network timeout, or malformed configuration cannot abort a multi-hour pentest assessment, providing resilient automation for security testing workflows.
Summary
- Promise.allSettled in
src/temporal/workflows.tsexecutes vulnerability and exploit pipelines in parallel without allowing individual failures to cancel the entire group. - Structured error classification in
src/temporal/activities.tsdistinguishes retryable errors (which trigger Temporal retries) from non-retryable errors (which result in graceful skips). - Conditional execution ensures that exploit agents only run when
checkExploitationQueuereturnsshouldExploit: true, preventing invalid or corrupted vulnerability data from causing downstream failures. - Result aggregation collects partial successes from healthy agents while logging specific failures, ensuring the workflow produces actionable reports even when some agents fail completely.
Frequently Asked Questions
What happens if one vulnerability agent crashes in Shannon?
If a single vulnerability agent crashes, Promise.allSettled in src/temporal/workflows.ts captures the rejection while allowing the other four parallel pipelines to continue executing. The workflow logs the specific error message but does not terminate, ensuring that findings from healthy agents are preserved in the final report.
How does Shannon decide whether to retry a failed activity or skip it?
Shannon's checkExploitationQueue function in src/temporal/activities.ts inspects the retryable property of caught errors. If error.retryable is true, the function re-throws the error to trigger Temporal's built-in retry mechanism. If false, it returns a skip decision (shouldExploit: false), allowing the workflow to bypass the failed agent without aborting the entire execution group.
Can Shannon produce a report if all exploit agents fail?
Yes. Shannon aggregates metrics from any successful vulnerability agents even if all corresponding exploit agents fail or are skipped. The reporting phase in the Temporal workflow processes whatever data was successfully gathered, producing a partial report that includes vulnerability findings without exploitation results. This ensures that valuable reconnaissance data is never lost due to downstream exploitation failures.
Where is the parallel execution logic implemented in the Shannon codebase?
The parallel execution logic resides primarily in src/temporal/workflows.ts, specifically around line 198 where Promise.allSettled orchestrates the five vulnerability-to-exploit pipelines. The supporting error classification logic that enables graceful failure handling is implemented in src/temporal/activities.ts within the checkExploitationQueue function starting at line 40.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →