How INFINI Gateway's Automatic Retry Mechanism Handles Transient Elasticsearch Failures
INFINI Gateway implements a configurable retry loop in proxy/output/elastic/reverseproxy.go that detects transient Elasticsearch failures, marks unhealthy nodes, and automatically retries requests on alternative hosts up to a user-defined maximum.
The automatic retry mechanism in INFINI Gateway provides resilient reverse-proxying to Elasticsearch clusters by intercepting transient network failures and overloaded node responses. According to the infinilabs/gateway source code, this functionality is implemented in the Elasticsearch output proxy component, which evaluates error patterns against a whitelist of transient conditions before attempting failovers to healthy nodes.
How the Automatic Retry Mechanism Detects Transient Failures
Identifying Recoverable Error Patterns
In proxy/output/elastic/reverseproxy.go (lines 45-73), the proxy maintains a failureMessage whitelist containing string patterns such as "connection refused", "timeout", "no such host", and HTTP 429 Too Many Requests. When a request returns an error matching these patterns, the proxy classifies the incident as a transient backend failure rather than a permanent application error, triggering the retry workflow.
Marking Nodes as Unhealthy
Upon detection, the gateway immediately reports the failure via elastic.GetOrInitHost(host, …).ReportFailure() (lines 80-84). This updates the internal health-monitoring state, ensuring the load balancer excludes the degraded node from subsequent selection rounds until the node recovers.
Retry Eligibility and Failover Logic
Configuration-Driven Retry Decisions
The retry logic consults several boolean flags defined in proxy/output/elastic/config.go to determine whether a specific request should be retried (lines 85-113 in reverseproxy.go):
- RetryOnBackendFailure: Master switch enabling automatic retries for network-level failures
- RetryReadonlyOnlyOnBackendFailure: Permits retries for idempotent GET/HEAD requests
- RetryWriteOpsOnBackendFailure: Controls retries for POST/PUT/PATCH/DELETE (carries risk of duplicate writes)
- RetryOnBackendBusy: Specific handling for HTTP 429 responses indicating an overloaded node
Host Exclusion and Failover
To prevent retry storms against the same failing node, the proxy maintains a hashset.Set named skippedHost (lines 89-96). When a retry is initiated, the failing host is added to this set, forcing the load balancer to select an alternative Elasticsearch node from the cluster for the next attempt.
Retry Limits, Timing, and Observability
Enforcing Retry Boundaries
A counter tracks retry attempts against the MaxRetryTimes configuration value (lines 121-129). When MaxRetryTimes is set to 0, retries are unlimited; otherwise, the proxy aborts the request and returns an error response to the client once the limit is exceeded.
Back-off Delays and Request Metadata
When RetryDelayInMs is configured and the retry targets the same host, the proxy inserts a sleep interval (lines 130-132) before reissuing the request, reducing pressure on struggling nodes. Each retry attaches diagnostic metadata: the request header RETRY_AT receives a timestamp, and the final response includes X-Retry-Times indicating the total attempt count (lines 134-139). Downstream services can inspect these headers to monitor cluster stability and retry efficiency.
Configuring the Automatic Retry Mechanism
The retry behavior is controlled through the ProxyConfig struct in proxy/output/elastic/config.go. Below is a sample gateway.yml configuration:
proxy:
elasticsearch: "my-es-cluster"
balancer: "round_robin"
max_retry_times: 5
retry_on_backend_failure: true
retry_readonly_on_backend_failure: true
retry_writes_on_backend_failure: false
retry_on_backend_busy: true
retry_delay_in_ms: 200
max_retry_times: Maximum attempts (0 = unlimited)retry_delay_in_ms: Millisecond pause between retries on the same host
Summary
The automatic retry mechanism in INFINI Gateway provides cluster resilience through the following workflow:
- Transient failure detection via
failureMessagewhitelist matching inreverseproxy.go(lines 45-73) - Health state management through
ReportFailure()calls that mark nodes unavailable - Method-specific retry controls distinguishing between safe read operations and potentially duplicative writes
- Host-level failover using
skippedHostsets to ensure distinct nodes are attempted - Configurable boundaries via
MaxRetryTimesand optionalRetryDelayInMsback-off intervals - Request tracing via
X-Retry-TimesandRETRY_ATheaders for operational observability
Frequently Asked Questions
What types of Elasticsearch failures trigger automatic retries?
Connection refused, timeout, DNS resolution failures ("no such host"), and HTTP 429 Too Many Requests responses trigger the retry loop. The proxy matches error strings against an internal failureMessage whitelist in proxy/output/elastic/reverseproxy.go (lines 45-73).
How does INFINI Gateway prevent retrying against the same failed node?
The proxy uses a hashset.Set named skippedHost to track failed nodes during the retry cycle (lines 89-96). When selecting a host for subsequent attempts, the load balancer excludes any node present in this set, ensuring failover to healthy instances.
Can I safely retry write operations through the gateway?
The RetryWriteOpsOnBackendFailure flag enables retries for POST/PUT/PATCH/DELETE requests, but this carries a risk of duplicate writes if the original request partially succeeded. For data safety, the default configuration only retries idempotent read operations (GET/HEAD) via RetryReadonlyOnlyOnBackendFailure.
How can I monitor retry activity in my application responses?
Inspect the X-Retry-Times response header, which indicates how many attempts were required before success. The proxy also adds a RETRY_AT timestamp header to outgoing requests during retry cycles, allowing downstream systems to track retry timing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →