How to Configure Automatic Retries for Steps in Probe
Probe supports automatic retries through the retry field in step definitions, using max_attempts, interval, and initial_delay parameters controlled by the StepRetry struct and executed via executeActionWithRetry in step.go.
The linyows/probe workflow engine allows you to configure automatic retries for steps that may fail due to transient issues. By adding a retry block to any step definition, you can specify exactly how many times Probe should attempt execution and how long to wait between tries. This guide explains the configuration schema, implementation details from the source code, and practical examples for handling flaky operations.
Understanding the Retry Configuration Structure
Probe defines retry behavior through the StepRetry struct located in repeat.go (lines 16-22). When you add a retry field to a step, you configure three key parameters that control the retry loop.
StepRetry Struct Fields
max_attempts(required): An integer ≥ 1 specifying the maximum number of execution attempts. This is enforced byStepRetry.Validateinrepeat.go(lines 23-31), which also checks against the global environment limit.interval(optional): The wait duration between consecutive attempts, specified as a Go duration string (e.g.,2s,1m).initial_delay(optional): The wait duration before the first attempt begins, allowing you to defer execution when necessary.
How Probe Executes Retries
When a step contains a retry block, the Step.Do method routes execution to executeActionWithRetry in step.go (lines 141-200). This function implements a controlled retry loop with the following algorithm:
- Initial Delay: If
initial_delayis configured, the runner sleeps for the specified duration before attempting execution. - Retry Loop: The system repeats up to
max_attempts, executing the action viaexecuteSingleActioneach iteration. - Success Check: If the action returns
ExitStatusSuccess(exit code 0), the loop terminates immediately. - Backoff: On failure or timeout, the runner sleeps for the configured
intervalbefore the next attempt. - Final Result: After exhausting all attempts, the function returns the last result and error.
When running with the --verbose flag, Probe logs detailed debug information showing the retry timeline and exit statuses for troubleshooting. The behavior is verified by the test suite in step_test.go, which confirms timeout handling and retry logic.
YAML Configuration Examples
Add the retry field directly under any step in your workflow definition. Here is a complete example checking a health endpoint:
jobs:
- name: Check Service
steps:
- name: Ping endpoint
uses: http
with:
url: https://example.com/health
get: /
retry:
max_attempts: 5
interval: 2s
initial_delay: 1s
test: res.code == 200
In this configuration:
- The step executes up to five times if previous attempts fail.
- Probe waits 2 seconds between consecutive attempts.
- The first attempt waits 1 second after the step starts.
Global Limits and Environment Controls
Probe enforces a safety ceiling on retries to prevent runaway executions. The PROBE_MAX_ATTEMPTS environment variable controls the maximum allowed value for max_attempts, defaulting to 10,000.
You can increase this limit without modifying source code:
export PROBE_MAX_ATTEMPTS=20000
probe workflow.yml
If a workflow specifies max_attempts exceeding this ceiling, StepRetry.Validate returns a clear configuration error during startup (see the validation logic in repeat.go, lines 23-31). The probe.go file handles top-level validation of these job configurations before execution begins.
Summary
- Configure retries by adding a
retryblock withmax_attempts,interval, andinitial_delayto any step definition. - The
StepRetrystruct inrepeat.godefines the schema, whileexecuteActionWithRetryinstep.goimplements the execution loop. - Probe respects the
PROBE_MAX_ATTEMPTSenvironment variable (default 10,000) as a global safety limit. - Use
--verboselogging to debug retry timelines and exit statuses.
Frequently Asked Questions
What is the maximum number of retries allowed in Probe?
Probe defaults to a maximum of 10,000 attempts per step, controlled by the PROBE_MAX_ATTEMPTS environment variable. You can increase this by exporting the variable before running your workflow, but any max_attempts value exceeding this ceiling triggers a validation error from StepRetry.Validate in repeat.go.
How does Probe determine if a step succeeded during retries?
The retry loop in step.go checks the exit status after each call to executeSingleAction. If the status equals ExitStatusSuccess (value 0), the loop breaks immediately and returns success. Any non-zero exit code or timeout triggers the next retry attempt after the configured interval.
Can I set different retry intervals for each step in a workflow?
Yes, each step defines its own independent retry configuration. The interval and initial_delay fields are scoped per step, allowing fine-grained control over timing for HTTP requests, database connections, or command executions within the same workflow.
Where does Probe validate the retry configuration?
Validation occurs in multiple layers. The StepRetry.Validate method in repeat.go checks that max_attempts meets minimum requirements and respects the global environment ceiling. Additionally, probe.go performs top-level validation of job and step configurations when loading the workflow file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →