How Poly Data Retries on Network Failures: A Deep Dive into the Resilience Strategy
Poly Data implements a multi-layered retry strategy with fixed back-offs for HTTP 500/429 errors, capped exponential back-off for GraphQL queries, and configurable retry limits for transient failures across poly_utils/utils.py, parallel_sync.py, and update_utils/update_markets.py.
Network reliability is critical when aggregating data from Polymarket's Gamma API and Goldsky's GraphQL endpoints. The warproxxx/poly_data repository embeds explicit retry logic directly into its data fetching layers to survive temporary network glitches and API throttling. Understanding how Poly Data retry on network failures works reveals a robust framework designed to keep data pipelines running despite intermittent service disruptions.
HTTP Status Code-Based Retry Logic
The codebase distinguishes between specific HTTP error codes to apply appropriate retry strategies. Each status code triggers a distinct back-off duration and retry limit, ensuring efficient use of resources while maximizing recovery chances.
Handling Server Errors (HTTP 500)
When the Polymarket API returns an HTTP 500 response during bulk market pagination, the system waits 5 seconds before retrying the same request. In update_utils/update_markets.py, this logic appears inside the batch processing loop:
if response.status_code == 500:
print("Server error (500) - retrying in 5 seconds...")
time.sleep(5)
continue
This simple loop structure allows unlimited retries for server errors, relying on the surrounding while/for loops to maintain the batch context. The code jumps back to the start of the batch loop after sleeping, attempting the request again without incrementing a retry counter.
Rate Limiting (HTTP 429)
HTTP 429 responses trigger longer back-off periods to respect API rate limits. Both the token-level fetch in poly_utils/utils.py and the bulk pagination in update_utils/update_markets.py implement 10-second sleeps when encountering rate limits:
elif response.status_code == 429:
print("Rate limited (429) - waiting 10 seconds...")
time.sleep(10)
continue
This strategy applies unlimited retries while the request remains rate-limited, preventing data loss during high-traffic periods.
Generic API Failures
For non-200, non-429 responses, Poly Data enforces a maximum of 3 attempts with shorter back-off intervals. In poly_utils/utils.py (lines 95-112), the retry counter increments before sleeping 2 seconds (token fetch) or 3 seconds (bulk fetch):
retry_count = 0
max_retries = 3
while retry_count < max_retries:
response = requests.get(..., timeout=30)
if response.status_code == 429:
time.sleep(10)
continue
elif response.status_code != 200:
retry_count += 1
time.sleep(2)
continue
Once retry_count reaches 3, the loop exits, allowing the application to handle the persistent failure gracefully.
Exponential Back-Off for GraphQL Queries
Goldsky GraphQL queries implement a more sophisticated retry mechanism using exponential back-off capped at 30 seconds. In parallel_sync.py, the goldsky_query function attempts the request up to 5 times before failing:
for attempt in range(5):
resp = session.post(QUERY_URL, json={'query': query}, timeout=30)
resp.raise_for_status()
data = resp.json()
...
except Exception as e:
wait = min(2 ** attempt, 30)
print(f" [retry {attempt+1}] {e} — waiting {wait}s")
time.sleep(wait)
This pattern waits 2, 4, 8, 16, and 30 seconds between successive attempts, preventing thundering herd problems while maximizing the chance of recovery from service-side throttling.
Network-Level Exception Handling
Beyond HTTP status codes, Poly Data catches low-level network failures using requests.exceptions.RequestException. In update_utils/update_markets.py (lines 64-70), a broad exception handler catches timeouts and connection errors:
except requests.exceptions.RequestException as e:
print(f"Error fetching batch: {e}")
time.sleep(5)
continue
This catch-all mechanism sleeps 5 seconds before retrying, ensuring that transient DNS failures, SSL handshake timeouts, or connection drops do not abort the entire data gathering run.
Implementation Examples
The retry logic operates transparently within high-level utility functions, requiring no manual intervention during normal operation.
Retrying Token-Level Market Fetches
When fetching missing market data for specific tokens, the fetch_missing_markets function automatically applies the 3-retry limit:
from poly_utils.utils import fetch_missing_markets
# Automatically retries up to 3 times on non-200 responses,
# waiting 2s between attempts and 10s on HTTP 429
fetch_missing_markets(missing_token_ids)
Querying Goldsky with Automatic Retries
Direct GraphQL queries benefit from exponential back-off without additional wrapper code:
from parallel_sync import goldsky_query
import requests
session = requests.Session()
where_clause = 'timestamp_gt: "1622505600", timestamp_lte: "1625097600"'
# Attempts up to 5 times with exponential back-off (2, 4, 8, 16, 30s)
events = goldsky_query(session, where_clause)
print(f"Fetched {len(events)} orderFilled events")
Bulk Market Pagination with Error Recovery
The bulk fetch utility handles server errors and rate limits while paginating through large datasets:
from update_utils.update_markets import fetch_all_markets
# Runs paginated fetch with automatic retry on HTTP 500 (5s wait),
# HTTP 429 (10s wait), and network exceptions (5s wait)
fetch_all_markets()
Summary
- Poly Data embeds retry logic directly into
poly_utils/utils.py,parallel_sync.py, andupdate_utils/update_markets.pyto handle network instability. - Fixed back-offs of 5 seconds (HTTP 500) and 10 seconds (HTTP 429) apply to REST API calls with unlimited retries.
- Capped retries (maximum 3 attempts) with 2-3 second delays protect against persistent non-200 errors.
- Exponential back-off (capped at 30 seconds) over 5 attempts protects GraphQL queries in
parallel_sync.py. - Network-level exceptions trigger 5-second retries via
requests.exceptions.RequestExceptionhandlers.
Frequently Asked Questions
What is the maximum number of retries for market data fetching?
Token-level fetches in poly_utils/utils.py enforce 3 retries for generic API failures, while Goldsky GraphQL queries in parallel_sync.py attempt 5 retries. HTTP 500 and 429 errors trigger unlimited retries with fixed back-off periods until the service recovers.
How does Poly Data handle GraphQL query failures differently from REST API failures?
GraphQL failures use exponential back-off (2^attempt seconds, capped at 30s) over 5 attempts, whereas REST API failures use fixed back-off periods (5s for HTTP 500, 10s for HTTP 429, 2-3s for other errors). This distinction reflects the different throttling behaviors of the Goldsky endpoint versus Polymarket's Gamma API.
What happens when Poly Data encounters a network timeout?
The system catches requests.exceptions.RequestException in update_utils/update_markets.py, prints the error details, sleeps 5 seconds, and retries the current batch. This catch-all handler ensures DNS failures, connection drops, and SSL timeouts do not terminate the data synchronization process.
Where is the retry logic configured in the Poly Data codebase?
Retry configurations are hardcoded within specific data fetching functions: poly_utils/utils.py (lines 95-112) for token-level retries with max_retries = 3, parallel_sync.py (lines 87-102) for GraphQL exponential back-off, and update_utils/update_markets.py (lines 64-88) for bulk pagination error handling. No centralized configuration file exists; each module implements context-appropriate retry strategies.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →