Configuring Retry Logic for Transient vs Permanent Errors in Shannon Temporal Activities

Shannon's Temporal activities use distinct retry policies defined in src/temporal/workflows.ts that classify AuthenticationError, PermissionError, and five other specific error types as permanent (non-retryable), while treating all other errors as transient with exponential back-off.

Shannon, an open-source data pipeline framework maintained by KeygraphHQ, orchestrates every pipeline step as a Temporal activity. The repository implements a sophisticated retry logic configuration that distinguishes between transient failures requiring automatic recovery and permanent errors demanding immediate human intervention.

Production vs Testing Retry Policies

Shannon provides two distinct retry policy objects to accommodate different operational contexts. Both are exported from src/temporal/workflows.ts and referenced when creating activity proxies.

Production Retry Policy (PRODUCTION_RETRY) is optimized for long-running pipelines that must survive temporary infrastructure hiccups. It configures aggressive retry behavior with extended intervals to outlast transient outages without manual intervention.

Testing Retry Policy (TESTING_RETRY) prioritizes rapid feedback during development and CI/CD pipelines. It reuses the same permanent error classifications but shortens retry intervals and reduces attempt counts to fail fast on genuine issues.

Permanent Error Types That Disable Retries

The nonRetryableErrorTypes array in src/temporal/workflows.ts (lines 44–59) explicitly lists seven error classifications that Shannon treats as permanent. When an activity throws an error matching any of these types, Temporal aborts execution immediately without retry attempts.

The permanent error types are:

  • AuthenticationError – Invalid API keys or expired credentials
  • PermissionError – Insufficient access rights to target resources
  • InvalidRequestError – Malformed API calls or schema violations
  • RequestTooLargeError – Payload size exceeding service limits
  • ConfigurationError – Invalid pipeline or activity configuration
  • InvalidTargetError – Non-existent or unreachable target endpoints
  • ExecutionLimitError – Hard limits exceeded (quota, budget, etc.)

How Shannon Classifies Errors

Error classification logic resides in src/error-handling.ts within the classifyErrorForTemporal function (lines 11–39). This utility inspects raw error messages and maps them to the specific types listed above, returning a tuple containing the error type and a boolean indicating retryability.

When an activity encounters an exception, it calls classifyErrorForTemporal to determine the appropriate error classification. The function uses pattern matching to identify authentication failures, permission denials, and other permanent conditions. Errors matching billing-related patterns or output-validation issues are marked as retryable, while unrecognized errors default to a generic TransientError classification with retryable: true.

Activities then wrap the original error using Temporal's ApplicationFailure.fromError(), passing the classified type and the retryable flag. This allows Temporal's retry machinery to consult the activity proxy's nonRetryableErrorTypes list and make the appropriate retry decision.

Configuring Retry Back-Off Behavior

The retry policies in src/temporal/workflows.ts define specific back-off parameters that control how Temporal schedules retry attempts for transient errors.

Production settings (lines 44–59):

  • Initial interval: 5 minutes
  • Maximum interval: 30 minutes
  • Back-off coefficient: 2 (doubles each attempt)
  • Maximum attempts: 50

Testing settings (lines 61–68):

  • Initial interval: 10 seconds
  • Maximum interval: 30 seconds
  • Back-off coefficient: 2
  • Maximum attempts: 5

These configurations ensure that production pipelines can survive extended service outages or rate-limiting periods while testing environments provide rapid feedback without wasting compute on prolonged retry cycles.

Implementation in Activity Proxies

Shannon implements the retry logic by passing the appropriate policy object when creating Temporal activity proxies. The following examples demonstrate how to configure activities for both production and testing contexts.

Production activity proxy with extended timeouts and aggressive retry policy:

import { proxyActivities } from '@temporalio/workflow';
import { PRODUCTION_RETRY } from './workflows';
import * as activities from './activities';

const acts = proxyActivities<typeof activities>({
  startToCloseTimeout: '2 hours',
  heartbeatTimeout: '60 minutes',
  retry: PRODUCTION_RETRY,  // 5min initial, 50 attempts, permanent errors abort immediately
});

Testing activity proxy for rapid iteration:

import { proxyActivities } from '@temporalio/workflow';
import { TESTING_RETRY } from './workflows';
import * as activities from './activities';

const testActs = proxyActivities<typeof activities>({
  startToCloseTimeout: '30 minutes',
  heartbeatTimeout: '30 minutes',
  retry: TESTING_RETRY,  // 10s initial, 5 attempts, same permanent error classification
});

Error classification within an activity:

import { ApplicationFailure } from '@temporalio/activity';
import { classifyErrorForTemporal } from '../error-handling';

try {
  // Activity implementation
  await processData();
} catch (err) {
  const { type, retryable } = classifyErrorForTemporal(err);
  
  // Wrap error so Temporal can apply retry policy logic
  throw ApplicationFailure.fromError(err, type, { nonRetryable: !retryable });
}

Summary

  • Shannon defines retry policies in src/temporal/workflows.ts with PRODUCTION_RETRY for long-running pipelines and TESTING_RETRY for rapid development cycles.
  • Seven specific error types—including AuthenticationError, PermissionError, and ConfigurationError—are classified as permanent and never retried.
  • All other errors default to transient status and retry with exponential back-off (5-minute initial intervals in production, 10 seconds in testing).
  • The classifyErrorForTemporal function in src/error-handling.ts maps raw exceptions to typed errors that Temporal's retry machinery evaluates against the nonRetryableErrorTypes list.

Frequently Asked Questions

What specific error types does Shannon consider permanent and non-retryable?

Shannon treats seven error classifications as permanent: AuthenticationError, PermissionError, InvalidRequestError, RequestTooLargeError, ConfigurationError, InvalidTargetError, and ExecutionLimitError. These are explicitly listed in the nonRetryableErrorTypes array within src/temporal/workflows.ts (lines 44–59), ensuring Temporal aborts execution immediately rather than wasting cycles on unrecoverable conditions like invalid credentials or malformed configurations.

How does Shannon's testing retry policy differ from production?

The TESTING_RETRY policy defined in src/temporal/workflows.ts (lines 61–68) uses significantly shorter intervals and fewer attempts than PRODUCTION_RETRY. Testing starts with a 10-second initial interval (versus 5 minutes in production), caps at 30 seconds (versus 30 minutes), and limits attempts to 5 (versus 50). Both policies share the identical nonRetryableErrorTypes list, ensuring permanent errors fail fast in both environments while transient errors resolve quickly during development.

Where is the error classification logic implemented in Shannon?

Error classification resides in src/error-handling.ts within the classifyErrorForTemporal function (lines 11–39). This utility inspects error messages and stack traces to determine whether an exception represents a permanent condition (like authentication failure) or a transient issue (like rate limiting). It returns a tuple containing the error type string and a boolean retryable flag, which activities then pass to Temporal's ApplicationFailure constructor to trigger the appropriate retry behavior.

How do I configure custom retry behavior for a specific Shannon activity?

To customize retry behavior, create a new retry policy object following the structure defined in src/temporal/workflows.ts and pass it to proxyActivities when defining your activity proxy. You can adjust initialInterval, maximumInterval, backoffCoefficient, maximumAttempts, and nonRetryableErrorTypes to suit specific requirements. For example, activities calling external APIs with strict rate limits might use shorter initial intervals but fewer maximum attempts, while database operations might tolerate longer back-off periods.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →