How to Resume Goldsky Data Scraping in Poly Data: Complete Guide
Poly Data uses a JSON cursor state file to remember the last timestamp and event ID, allowing the scraper to resume exactly where it left off without re-downloading historical data.
The warproxxx/poly_data repository provides a robust Goldsky subgraph scraper that handles interruptions gracefully. Instead of fetching the entire dataset on every run, the system maintains a persistent cursor that tracks progress through the orderFilled events, ensuring efficient incremental updates.
How the Resume Mechanism Works
The scraper implements a stateful pagination system that persists progress to disk after every successful batch. This design protects against network failures, script interruptions, or system crashes.
Step 1: Initialize the Directory
Before fetching data, the script ensures the goldsky directory exists to store both the output CSV and the cursor state file.
In update_utils/update_goldsky.py (lines 18‑20), the scraper checks:
if not os.path.isdir('goldsky'):
os.makedirs('goldsky', exist_ok=True)
Step 2: Load the Cursor State
The get_latest_cursor() function (lines 33‑44) attempts to read goldsky/cursor_state.json. If present, it returns the stored last_timestamp, last_id, and optional sticky_timestamp.
If no cursor file exists, the function falls back to reading the last line of goldsky/orderFilled.csv using the Unix tail command for speed (lines 65‑81), with a pandas fallback for compatibility:
result = subprocess.run(['tail', '-n', '1', csv_path], capture_output=True, text=True)
This timestamp becomes the starting point for the next query.
Step 3: Handle Sticky Timestamps
When many events share the same timestamp (common in high-frequency trading data), standard timestamp-based pagination would skip records. The scraper detects this condition (lines 80‑86) and enters "sticky" mode:
if batch_first_timestamp == batch_last_timestamp:
sticky_timestamp = batch_first_timestamp
In this mode, the cursor keeps the timestamp fixed and paginates by id instead, ensuring no events are missed within the same second.
Step 4: Execute the Query Loop
Each iteration builds a GraphQL where clause based on the current cursor state (lines 20‑26). Normal mode uses timestamp_gt, while sticky mode uses both timestamp and id_gt:
if sticky_timestamp is not None:
where_clause = f'{{timestamp: "{sticky_timestamp}", id_gt: "{last_id}"}}'
else:
where_clause = f'{{timestamp_gt: "{last_timestamp}"}}'
Step 5: Persist State After Each Batch
After writing a batch to CSV, save_cursor() (lines 21‑23) immediately writes the latest position to disk:
save_cursor(last_timestamp, last_id, sticky_timestamp)
This atomic write ensures that an interrupted run can resume precisely from the last committed record.
Step 6: Clean Up on Completion
When the scraper reaches the end of available data, it removes the cursor file (lines 27‑30), signaling a fresh start for the next full run:
if os.path.isfile(CURSOR_FILE):
os.remove(CURSOR_FILE)
Core Files and Responsibilities
Understanding the module structure helps with debugging and custom integration.
update_utils/update_goldsky.py– Contains the mainscrape()function, cursor logic (get_latest_cursor(),save_cursor()), and the GraphQL query loop.parallel_sync.py– Provides the low‑levelgoldsky_query()helper function used by the scraper to execute GraphQL requests.update_all.py– The convenience entry point that orchestrates market updates, Goldsky scraping viaupdate_goldsky(), and live data processing.update_utils/process_live.py– Demonstrates how downstream modules consume the generatedorderFilled.csvfile.
Practical Code Examples
Run a Full Update (Auto-Resume)
Execute the main entry point to continue from the existing cursor:
python update_all.py
The script prints "Updating goldsky" and invokes update_goldsky(), which automatically detects and loads any existing cursor state.
Resume Only the Goldsky Scraper
For targeted updates without running the full pipeline:
python -c "from update_utils.update_goldsky import update_goldsky; update_goldsky()"
If goldsky/cursor_state.json exists, scraping continues from that point; otherwise, it starts from the end of the CSV or from timestamp zero.
Inspect or Reset the Cursor Manually
Force a full re-scrape by removing the cursor file, or check current progress:
import json
import os
CURSOR = "goldsky/cursor_state.json"
# Display current cursor state
if os.path.isfile(CURSOR):
with open(CURSOR, 'r') as f:
print(json.load(f))
else:
print("No cursor file found. Scraper will start from CSV tail or timestamp 0.")
# Uncomment to force restart:
# os.remove(CURSOR)
Verify Sticky Mode Status
Check if the scraper is currently paginating within a single timestamp:
from update_utils.update_goldsky import get_latest_cursor
timestamp, last_id, sticky = get_latest_cursor()
print(f"Timestamp: {timestamp}, LastID: {last_id}, Sticky: {sticky}")
When sticky is not None, the next query will constrain both timestamp and id_gt to avoid skipping events.
Summary
- Poly Data avoids redundant downloads by maintaining a
cursor_state.jsonfile that records the last successful fetch position. - The resume logic lives in
update_utils/update_goldsky.py, specifically withinget_latest_cursor()andsave_cursor(). - Sticky timestamp handling prevents data loss when multiple events occur in the same second by switching from timestamp-based to ID-based pagination.
- The system falls back to reading the CSV tail if no cursor exists, ensuring continuity even if the JSON state is manually deleted.
- State is persisted before each network request, making the scraper resilient to crashes and interruptions.
Frequently Asked Questions
What happens if the cursor file becomes corrupted?
If goldsky/cursor_state.json contains invalid JSON or is unreadable, get_latest_cursor() catches the exception and falls back to reading the last line of goldsky/orderFilled.csv. If the CSV is also empty or missing, the scraper starts from timestamp zero, effectively performing a full refresh.
How does the sticky timestamp feature prevent duplicate or missed records?
When a batch of fetched records all share the same timestamp, simply incrementing the timestamp would skip the remaining events from that second. The scraper detects this condition and sets sticky_timestamp, keeping the timestamp fixed while paginating by id_gt. This ensures every event is captured even when thousands occur within the same block timestamp.
Can I run the Goldsky scraper independently of the main update script?
Yes. While update_all.py provides the standard orchestration, you can import and run update_goldsky() directly from update_utils/update_goldsky.py. This is useful for debugging, testing specific date ranges, or running the scraper on a different schedule than the market updates.
Where does the scraper store the actual event data?
Raw order-filled events are appended to goldsky/orderFilled.csv (created at runtime). The cursor state is stored separately in goldsky/cursor_state.json. This separation allows you to archive or move the CSV without affecting the resume capability, provided you preserve the JSON cursor file or accept a fallback to the CSV tail on the next run.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →