How to Resume Market Data Collection in Poly Data: A Complete Guide
You can resume market data collection in Poly Data by simply re-running the update_markets script, which automatically detects existing CSV records and continues fetching from the last offset using idempotent pagination logic.
The warproxxx/poly_data repository provides a robust Python utility for collecting Polymarket data that handles interruptions gracefully. Whether your script crashes, your connection drops, or you manually stop the process, the built-in resume functionality ensures you never lose progress or fetch duplicate records.
How the Resume Mechanism Works in Poly Data
The core resumption logic resides in update_utils/update_markets.py. The update_markets() function implements an idempotent design that treats every run as a potential continuation of a previous session.
Detecting Existing Records
When the script starts, it calls the count_csv_lines() helper to inspect the target CSV file (default markets.csv). This function returns the number of data rows while ignoring the header line.
current_offset = count_csv_lines(csv_filename)
file_exists = os.path.exists(csv_filename) and current_offset > 0
If the file contains existing records, the script uses this row count as the starting offset for API requests. According to the source code in update_utils/update_markets.py (lines 7-16), this detection happens before any network calls, ensuring zero unnecessary API usage when resuming.
Choosing the Correct File Mode
Based on the existence check, the script selects the appropriate file mode to prevent header duplication:
- New collection: Mode
'w'writes the CSV header to a fresh file - Resume operation: Mode
'a'appends new data without rewriting the header
if file_exists:
print(f"Found {current_offset} existing records. Resuming from offset {current_offset}")
mode = 'a'
else:
print(f"Creating new CSV file: {csv_filename}")
mode = 'w'
This logic in update_utils/update_markets.py (lines 42-48) guarantees that interrupting and restarting the collection never results in malformed CSV files with multiple headers.
Paginated API Fetching with Offsets
The Polymarket API supports pagination via limit and offset parameters. The script constructs requests using the current offset calculated from your existing CSV:
params = {
'order': 'createdAt',
'ascending': 'true',
'limit': batch_size,
'offset': current_offset
}
After each batch processes successfully, the script increments the offset by the actual number of records written:
current_offset += batch_count
As implemented in update_utils/update_markets.py (lines 55-60 and 150-156), this approach ensures that when resuming, the first API request begins exactly where the previous run stopped, fetching only unfetched markets.
Running the Market Data Collection
You have two entry points for initiating or resuming collection. The primary method uses update_all.py, which orchestrates the complete data pipeline:
python -m update_all
For market-specific operations, import and call the utility directly:
python -c "from update_utils.update_markets import update_markets; update_markets()"
Both methods respect the resume logic automatically. The script creates markets.csv in your current working directory unless you specify a custom path via the csv_filename parameter.
Error Handling During Resumption
The collection script includes resilient error handling that preserves partial progress during network interruptions. According to update_utils/update_markets.py (lines 64-88), the system handles specific HTTP status codes with appropriate delays:
- HTTP 500: Retries after 5 seconds
- HTTP 429 (rate limiting): Waits 10 seconds before retry
- Other non-200 responses: Retries after 3 seconds
Network exceptions are caught and retried without corrupting the CSV file, allowing you to safely interrupt the script at any moment.
Summary
- Automatic detection: The script counts existing CSV rows to determine the resume offset in
count_csv_lines() - Safe appending: Uses file mode
'a'when resuming to prevent duplicate headers - Offset-based pagination: Continues API calls from the last fetched market using
current_offset - Fault tolerance: Implements specific retry logic for HTTP errors and network exceptions
- Simple execution: Re-run
update_markets()orupdate_all.pyto resume without manual configuration
Frequently Asked Questions
Does Poly Data duplicate data when resuming?
No. The update_markets function in update_utils/update_markets.py uses offset-based pagination that begins at the row count of your existing CSV. Since the API returns markets ordered by createdAt, and the offset accounts for already-collected records, resuming fetches only new data without duplicates.
What happens if I delete the CSV file mid-collection?
If you delete markets.csv or your custom filename between runs, the script detects the missing file and automatically starts a fresh collection from offset zero. The file_exists check relies on os.path.exists(), so removing the file effectively resets the collection progress.
How does Poly Data handle rate limits during resumption?
When the Polymarket API returns HTTP 429 (rate limit), the script pauses execution for 10 seconds before retrying the same request. This retry logic preserves your current offset and does not increment the counter until successful, ensuring no markets are skipped during rate-limited resumption operations.
Can I change the batch size when resuming?
Yes. You can pass a different batch_size parameter when calling update_markets() regardless of previous runs. The offset calculation depends only on the number of existing rows in the CSV, not the batch size used during previous fetches. However, consistency in batch size is recommended for predictable API behavior.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →